Detects model-stealing by query shape, not volume — so your best customer doesn't get banned and a patient attacker doesn't get through.
Volume-based rate limiting fails in both directions: block on volume and you throttle the customer paying you most; allow on volume and a patient extractor distills your model well under any threshold. The signal was never how much they query.
$ mextract compare
stream score coverage boundary repeats verdict
extraction 0.65 0.50 0.57 0.00 EXTRACTION
power-user 0.24 0.12 0.00 0.33 power user
Identical 400-query volume. Opposite verdicts. That's the whole thesis.
An extractor is trying to learn a function, so its traffic has a characteristic shape:
| Signal | Extractor | Legitimate power user |
|---|---|---|
| Coverage of the input space | sweeps it (learning everywhere) | clustered on their own business data |
| Boundary focus | probes low-confidence regions (max information) | boundary-agnostic; asks about real cases |
| Repeat rate | ~0% — a repeat teaches it nothing | high — real workloads re-query the same inputs |
| Hour entropy | uniform round the clock | diurnal, business hours |
| NN uniformity | evenly spaced, grid-like | organically clustered |
Volume appears nowhere in the score. It's used only as a confidence qualifier: below 50 queries the tool returns insufficient-data and no verdict, because a shape judgment on 20 calls isn't trustworthy.
Being "below threshold" isn't good enough — an operator needs to know a heavy user is legitimately heavy:
$ mextract assess --simulate power-user
✓ 33% repeated queries — consistent with a real workload hitting the same inputs
✓ queries are clustered, not sweeping the space
✓ clear diurnal pattern (business hours)
✓ queries are not boundary-seeking
→ heavy but legitimate-shaped usage: do NOT throttle
test_volume_alone_does_not_raise_score asserts that quadrupling a legitimate user's volume doesn't push them over the line.
git clone https://github.com/vinzabe/model-extraction-detector && cd model-extraction-detector
python -m pip install -e ".[dev]"
mextract compare # the table above
mextract assess --simulate extraction # exit 2
mextract assess --simulate power-user # exit 0
mextract assess --stream queries.json --json # your own trafficQuery stream is a JSON array of {"features": [...], "confidence": 0.87, "hour": 14}. Exit codes: 0 no extraction signature (including recognised power users), 2 extraction suspected, 1 error.
Recommendations are graduated — the tool suggests watermarking responses or a query budget rather than a ban, because the cost of being wrong about a customer is high.
python -m pip install -e ".[dev]"
pytest --cov=mextract # 20 tests, ~95% coverage
mypy --strict src/mextract # clean (3.12 target for numpy stubs)
ruff check src tests # cleanMIT © vinzabe