Every project fixes bugs, but a fix only stays fixed if some test notices when it goes away. Antibody puts a project's past bug fixes back, one at a time, on today's code and runs the project's own tests; if no test notices a fix being taken out, that bug can come back unnoticed.
Before: 13 of 20 caught. After: 18 of 20 caught.
Numbers come from audits/sqlparse/ledger.json.
Live ledger page: https://antibody-ledger.vercel.app/sqlparse/
Requires Linux (tested on Ubuntu), git, Python 3.12, and uv.
The runner imports Unix-only modules (fcntl, resource), so on Windows, run it inside WSL.
With IBM Bob: switch Bob to the Antibody mode and ask it to audit a repository (the recurrence-audit skill handles the full workflow).
By hand: run each step from the repository root:
python3 .bob/skills/recurrence-audit/scripts/antibody.py <step>
Steps, in order: setup, candidates, run, probe, prove, ledger, gate.
CI gate: .github/workflows/antibody.yml fails the build while a fix is still exposed, unless antibody-accepted.yaml lists it with a one-sentence reason for the acceptance.
- Rows with no data are unknown, never safe.
- "Caught" means a test noticed the regression, not that the test is perfect or complete.
- Antibody works with Python and pytest projects only.
- Antibody installs the audited project and runs its tests, so it runs that project's code: audit projects you would run anyway, or run it in CI or a container.
- Bob reads the project's issue and pull-request threads, which anyone can write. No status comes from Bob's word (the runner proves each one), but when auditing a project you don't trust, approve Bob's commands one at a time instead of auto-approving them.
Before the hackathon started we ran our own script — no Bob — against 18 Python libraries and published the results at https://github.com/aryangorde6/antibody-census. The headline: 479 bug fixes undone in 18 Python libraries, and 139 times every test still passed.
Bob handles the parts that need judgment:
- Reading the bug report and pull-request thread for each candidate fix, and deciding whether it was a real behavioural bug.
- Starting one subagent per exposed fix to write and prove a regression test.
- Writing the plain-English explanation for each row in the ledger.
The runner handles the parts that need determinism:
- Reversing each fix and running the project's own test suite.
- Re-checking every proof file to confirm the test still catches the bug.
Bob session exports are in bob_sessions/.
The last task (07) was Bob reviewing the finished repository read-only, on a second team member's account; each finding and what we did with it is in bob_sessions/07-review/triage.md.
See BOB.md for a detailed map of every file Bob created or edited.
| Source | Licence / note |
|---|---|
| Each audited project's GitHub repository (git history, and the issue and pull-request pages its fixes link to) | Project's own licence |
| GitHub Advisory Database | CC-BY 4.0 |
| PyPI | Used to install test dependencies |
Saved threads leave out author names, and @handles and email addresses are removed from their text; thread text stays on the machine that ran the audit, out of the repository.
MIT — see LICENSE.