Add GigaCheck-Detector-Multi span coverage submission - #204
Chandrakanth10 wants to merge 2 commits into
Conversation
|
Eval run succeeded! Link to run: link Here are the results of the submission(s): GigaCheck-Detector-Multi (span coverage aggregation)Release date: 2026-09-15 I've committed detailed results of this detector's performance on the test set to this PR. On the RAID dataset as a whole (aggregated across all generation models, domains, decoding strategies, repetition penalties, and adversarial attacks), it achieved an AUROC of 80.45 and a TPR of 54.54% at FPR=5% and 33.83% at FPR=1%. If all looks well, a maintainer will come by soon to merge this PR and your entry/entries will appear on the leaderboard. If you need to make any changes, feel free to push new commits to this PR. Thanks for submitting to RAID! |
This submits an independent full-test evaluation of the existing GigaCheck-Detector-Multi span detector, using span coverage to produce one ranking score per document. Model and checkpoint attribution remain with the original authors; this is not a new trained checkpoint or an author-endorsed submission.
9fabe1503f99cc8a6f639996f9475c3ec19bc067.Original text and whitespace are preserved. Documents are processed in 900-content-token windows with 120-token overlap, batch size one, BF16 backbone, FP32 span head, eager attention, and no confidence cutoff. Raw span endpoints are floored/ceiled and clipped to their window before mapping back to original character positions. Each character receives the maximum confidence of any covering span across all windows, or zero if uncovered; the document score is the float64 mean across all original characters. Higher means more AI evidence. This score is not calibrated authorship confidence or the proportion of AI-written words.
No training, fine-tuning, ensembling, threshold fitting, or access to hidden test labels was used for this submission. Upstream training provenance is not independently asserted.
date_releaseddenotes the public release of this aggregation configuration, not the original checkpoint release date. Versions, source revisions, dataset hashes, and the full aggregation definition are included inadditional_metadata.Validation completed locally for every expected ID, original input hash, tokenizer window, raw-span-derived score, checkpoint hash, and exported score: zero missing IDs, zero duplicate IDs, and zero score differences. All 61 pilot reference responses also matched exactly. The copied submission file was checked again against the validated export.
Prediction file SHA256:
203c3aebddff247fbc8bdddc1bf056b73517a1f5b523bece7ce9f038683089a6.Only
metadata.jsonandpredictions.jsonare included. Noresults.jsonor claimed benchmark accuracy is supplied; official scoring is left to the RAID evaluation workflow.