5707 stars · MIT — read from /blob/master/LICENSE, plain unmodified text, 'Copyright (c) 2024-2025 Chetan Jain' · no GitHub releases at all, ever — a determination, not a gap; PyPI botasaurus 4.0.97 published 2026-01-06, so the package is about eight months behind the code · Track this in Scout
A Python framework for collecting data from web pages that are trying to block automated visitors.
▶Repo detailsthe review · specs · pros & cons · install
What it is
A Python framework at 5,707 stars, created in May 2023, that wraps browser control and plain web requests behind short decorators. It names several commercial blocking products it aims to get past, and it can package a finished collector as a desktop application for Windows, macOS and Linux.
What it is good for. Someone who already writes collectors in Python and keeps getting refused by the site they need. The framework part is the real value: caching between runs, running many pages at once, and human-looking mouse movement are all supplied rather than written by hand.
- MIT licensed, read from the licence file itself and completely plain, so there is no paid key and no restriction on commercial use.
- Caching, parallel collection and browser control come as one piece, which is a large amount of fiddly work you do not have to write.
- It is alive. Code landed on 26 July 2026, which took some proving — see the appendix.
- It publishes no releases on GitHub at all. That is a decision, not an oversight, so there is no tag to pin and no changelog there; the version to follow is the published Python package, and that has not moved since 6 January 2026. The code is therefore about eight months ahead of what an ordinary install gives you.
- Nothing here is usable without writing Python. There is no window and no wizard, and the "desktop application" is a thing to be assembled rather than a thing to be downloaded.
- One maintainer is named in the licence file, and there are 52 open issues against 5 open pull requests. The documentation lives on a separate commercial website rather than in the repository, so it can move or change terms independently of the code.
- Its selling point is getting past named commercial blocking products, which usually means going against the terms of the site being read. That is a decision to make deliberately.
- scrapy/scrapy
The standard Python framework for collecting pages at scale, doing the same fetching and parsing job, but it ships no blocking-evasion and no browser.
Track this in Scout
apify/crawleeThe same idea of a framework with the awkward parts included, with proxy rotation and real browsers, but written for Node.js and TypeScript and backed by a company selling a hosted service.
Track this in Scout- ultrafunkamsterdam/undetected-chromedriver
Does only the evasion half, as a patched browser driver with no framework around it, and its GPL-3.0 is a much heavier obligation than MIT.
Track this in Scout
python3 -m venv venv && source venv/bin/activate python -m pip install --upgrade botasaurus



