Documents and knowledge · Edition No. 46 · 10 Oct 2026

py-pdf/pypdf

Joins, splits, locks and reads PDF files from a few lines of Python, with no service involved and nothing uploaded anywhere.

← Documents and knowledgeRead the whole edition →

10,250 stars · BSD-3-Clause read from /blob/main/LICENSE, holders FILLED IN (Mathieu Fenniak 2006-2008, with contributions by Ashish Kulkarni and Steve Witham), no added conditions; confirmed by pyproject.toml at the tag and by the package page. · 6.20.0 (2026-10-09), settled three ways: the release title itself carries the year, the package page says 'Released: Oct 9, 2026', and CHANGELOG.md read AT THE TAG says 2026-10-09. · Track this in Scout

Joins, splits, locks and reads PDF files from a few lines of Python, with no service involved and nothing uploaded anywhere.

▶Repo detailsthe review · specs · pros & cons · install

What it does

pypdf reads a PDF file, gives you its pages as objects, and writes a new file out. It merges several documents into one, pulls a page range into its own file, rotates and crops pages, and copies bookmarks, named destinations, annotations and viewer settings across. It extracts text, with an optional "layout mode" that tries to keep columns in the order a person would read them. It opens password-protected files and can encrypt the ones it writes — the basic install handles the older RC4 scheme, and the stronger AES needs the extra pypdf[crypto] install. It decodes the compression formats PDFs use, including Brotli, which arrived in version 6.20.0 on 9 October 2026. It is written in pure Python, so a plain install pulls in almost nothing else. It does not draw pages as images, so it is not a viewer and cannot make thumbnails. It does not do character recognition on scans, and it does not find tables or work out page layout — that is a different tool's job.

Why it matters

Who it suits. Anyone who handles PDFs in bulk: merging statements, splitting a scanned batch into one file per invoice, stripping passwords off files a supplier insists on protecting, or pulling text out of a hundred reports so something else can search them. It is also the right choice when the documents must not leave the building, because there is no service involved. Skip it if you want to see the pages — it has no interface — and skip it if the PDFs are photographs of paper, because it reads text that is already text.

What people say. Two independent distributions carry it and both name a maintainer and a date, which is the most useful outside signal a library like this gets. The Arch Linux package page, read 10 October 2026, lists python-pypdf 6.20.0-1 packaged by Caleb Maclennan and last updated 9 October 2026 — the same day as the upstream release. The Debian package tracker, also read 10 October 2026, records 6.19.0-1 accepted into unstable on 30 September 2026 by Pieter Lenaerts and migrated to testing on 5 October 2026. Neither sells anything. The one security write-up found, Wiz on CVE-2026-33123 published 20 March 2026, describes a crafted PDF driving excessive processor and memory use in versions before 6.9.1 — and Wiz sells a commercial cloud-security product, so its vulnerability database is marketing as well as information, although the facts check out against the project's own advisory list. No authored, dated, non-vendor technical review of pypdf exists that we could find, and that is the honest finding.

Verdict. The best five minutes in this edition. Install it, point it at a folder of PDFs, and it will do in one afternoon what a paid tool charges monthly for. The one thing to know is the security pattern: the project published ten advisories between 26 May and 23 June 2026, all of the same kind — a malformed PDF making it loop forever or eat memory — and it keeps shipping limits to stop them. So wrap it in a time limit and a memory cap before feeding it files from strangers, and stay on the current version. If you need to turn pages into pictures, pymupdf/PyMuPDF does that and was Edition 37's entry, but it is AGPL-3.0 with a paid commercial option rather than BSD. If you want tables out of a PDF, jsvine/pdfplumber is the tool for that job.

Stars10,250
LicenceBSD-3-Clause read from /blob/main/LICENSE, holders FILLED IN (Mathieu Fenniak 2006-2008, with contributions by Ashish Kulkarni and Steve Witham), no added conditions; confirmed by pyproject.toml at the tag and by the package page.
Latest6.20.0 (2026-10-09), settled three ways: the release title itself carries the year, the package page says 'Released: Oct 9, 2026', and CHANGELOG.md read AT THE TAG says 2026-10-09.
Good
  • BSD-3-Clause read from the file with the real copyright holders named, no added conditions, and a plain install that pulls in essentially nothing else.
  • Genuinely current: version 6.20.0 released 9 October 2026, confirmed by the project's own changelog, its package page and two independent Linux distributions.
  • It publishes actual hardening limits with numbers — 10,000 objects before it gives up recovering a damaged file, 500,000 objects while copying, a maximum object number of 1,000,000 — all of which can be turned off if you need to.
Watch for
  • Ten published advisories in four weeks of mid-2026, every one of them a malformed file causing an endless loop or runaway memory use. Current versions are patched, but the pattern says: do not hand it untrusted files without a timeout.
  • The stronger AES encryption, image extraction and right-to-left text all need extra installs, and the [full] option pulls in a component under the LGPL, which is a stricter licence than the library's own BSD. The documented route for one image format tells you to install jbig2dec, which is GPL3.
  • It is a library. There is nothing to open, nothing to click, and no command-line tool as the main way in.
Similar repositories
Install
pip install pypdf
pip install pypdf[crypto]
pip install pypdf[image]
Screenshots
py-pdf/pypdf: GitHub preview card

Get the next edition in your inbox

A dozen repositories, opened and checked. The licence read, the last release dated, and the ones that did not make it named with the reason. It is the half most lists leave out.

No tracking pixels. One click to leave. The archive stays free either way.

We use your address to send the edition and nothing else. Confirm by email, leave in one click. How we handle it.