Documents and knowledge · Edition No. 47 · 11 Oct 2026

jsvine/pdfplumber

Pulls text, tables and the exact position of every object out of a PDF, and draws the table it detected so a bad result can be fixed.

← Documents and knowledgeRead the whole edition →

10,700 stars · MIT, read from /blob/stable/LICENSE.txt, plain and unmodified, 'Copyright (c) 2015, Jeremy Singer-Vine' · v0.11.10 (2026-06-15), read from /releases/latest; the year was absent and was settled against pypi.org/project/pdfplumber/, which dates 0.11.10 to 15 June 2026 · Track this in Scout

Pulls text, tables and the exact position of every object out of a PDF, and draws the table it detected so a bad result can be fixed.

▶Repo detailsthe review · specs · pros & cons · install

What it does

You give it a file path, or bytes, or a file object, and a password if the file has one. It exposes every object on each page as a list you can read: characters, lines, rectangles, curves, images, annotations and links, each with its coordinates. On top of that it extracts text, words and tables. Table finding has four strategies in each direction — ruled lines, strict lines, text alignment, or boundaries you give it yourself — and twenty settings to tune them. There is also a command-line mode that writes CSV, JSON or plain text. The debugging view renders a page as an image at 72 dots per inch and draws the lines, intersections and table boundaries it detected on top, which is how you work out why a table came out wrong. What it does not do is stated plainly in the README: it "works best on machine-generated, rather than scanned, PDFs", it has no text recognition for scans, it does not create or modify PDFs at all, it does not rebuild the content of images, and it has no interface for form fields.

Why it matters

Who it suits. Anyone who regularly gets data as a PDF and types it out again. Bank statements, invoices, government reports and published tables are the usual cases, and this is the tool that gets them into rows and columns. It also suits anyone debugging a table extraction that another library got wrong, because the picture it draws shows you what the computer actually saw. Skip it if your documents are scans or photographs, because there is no text recognition here at all, and skip it if you need to write PDFs rather than read them.

Verdict. Worth twenty minutes today on a document you have already given up on once. pip install pdfplumber and four lines of Python gets you the first table, and the debugging image is what makes the second one work. Three honest limits. It is slower than the alternatives, and the README says so itself, naming pymupdf/PyMuPDF as substantially faster. On a large document its per-page caching "can use a lot of memory", released by closing each page. And several features, including layout-preserving text extraction and search, are marked experimental. For scans you need text recognition first, and Edition 27 covered ocrmypdf/OCRmyPDF, which makes a scan searchable and then hands it to a tool like this one. A note on checking: the usual mirror this report uses for code dates refused this repository on three attempts, so currency here rests on its release of 15 June 2026, confirmed at the package registry, and on a test list that includes Python 3.14.

Stars10,700
LicenceMIT, read from /blob/stable/LICENSE.txt, plain and unmodified, 'Copyright (c) 2015, Jeremy Singer-Vine'
Latestv0.11.10 (2026-06-15), read from /releases/latest; the year was absent and was settled against pypi.org/project/pdfplumber/, which dates 0.11.10 to 15 June 2026
Good
  • It draws a picture of the table it detected, which turns a failed extraction into a fixable problem.
  • Four table-finding strategies and twenty settings, so an awkward layout is usually reachable.
  • MIT licence, read from the file, plain, and a command-line mode for when you do not want to write code.
Watch for
  • No text recognition. A scanned document gets you nothing until you run it through something else first.
  • Slower than the alternatives, which the README states itself, and page caching can use a lot of memory on big files.
  • Version 0.11.10, before 1.0, with several features marked experimental, and the table system changed in a breaking way at 0.5.0.
Similar repositories
Install
python3 -m venv venv
source venv/bin/activate
pip install pdfplumber
# Command-line use
pdfplumber statement.pdf --format csv --pages 1-3 > rows.csv

Get the next edition in your inbox

A dozen repositories, opened and checked. The licence read, the last release dated, and the ones that did not make it named with the reason. It is the half most lists leave out.

No tracking pixels. One click to leave. The archive stays free either way.

We use your address to send the edition and nothing else. Confirm by email, leave in one click. How we handle it.