Skip to content

Repository files navigation

model-extraction-detector

Detects model-stealing by query shape, not volume — so your best customer doesn't get banned and a patient attacker doesn't get through.

Volume-based rate limiting fails in both directions: block on volume and you throttle the customer paying you most; allow on volume and a patient extractor distills your model well under any threshold. The signal was never how much they query.

$ mextract compare
stream         score  coverage  boundary  repeats       verdict
extraction      0.65      0.50      0.57     0.00    EXTRACTION
power-user      0.24      0.12      0.00     0.33    power user

Identical 400-query volume. Opposite verdicts. That's the whole thesis.

What actually distinguishes them

An extractor is trying to learn a function, so its traffic has a characteristic shape:

Signal Extractor Legitimate power user
Coverage of the input space sweeps it (learning everywhere) clustered on their own business data
Boundary focus probes low-confidence regions (max information) boundary-agnostic; asks about real cases
Repeat rate ~0% — a repeat teaches it nothing high — real workloads re-query the same inputs
Hour entropy uniform round the clock diurnal, business hours
NN uniformity evenly spaced, grid-like organically clustered

Volume appears nowhere in the score. It's used only as a confidence qualifier: below 50 queries the tool returns insufficient-data and no verdict, because a shape judgment on 20 calls isn't trustworthy.

It names the power user explicitly

Being "below threshold" isn't good enough — an operator needs to know a heavy user is legitimately heavy:

$ mextract assess --simulate power-user
  ✓ 33% repeated queries — consistent with a real workload hitting the same inputs
  ✓ queries are clustered, not sweeping the space
  ✓ clear diurnal pattern (business hours)
  ✓ queries are not boundary-seeking

  → heavy but legitimate-shaped usage: do NOT throttle

test_volume_alone_does_not_raise_score asserts that quadrupling a legitimate user's volume doesn't push them over the line.

Quickstart (60 seconds)

git clone https://github.com/vinzabe/model-extraction-detector && cd model-extraction-detector
python -m pip install -e ".[dev]"

mextract compare                                    # the table above
mextract assess --simulate extraction               # exit 2
mextract assess --simulate power-user               # exit 0
mextract assess --stream queries.json --json        # your own traffic

Query stream is a JSON array of {"features": [...], "confidence": 0.87, "hour": 14}. Exit codes: 0 no extraction signature (including recognised power users), 2 extraction suspected, 1 error.

Recommendations are graduated — the tool suggests watermarking responses or a query budget rather than a ban, because the cost of being wrong about a customer is high.

Development

python -m pip install -e ".[dev]"
pytest --cov=mextract      # 20 tests, ~95% coverage
mypy --strict src/mextract # clean (3.12 target for numpy stubs)
ruff check src tests       # clean

License

MIT © vinzabe

About

Detects model-extraction campaigns by query SHAPE not volume: separates boundary-probing distillation from a heavy legitimate user, so the power user is identified and not throttled.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages