feat(eval): add shared CTC ASR support for facebook/mms-1b-all - #1335
feat(eval): add shared CTC ASR support for facebook/mms-1b-all#1335ssss141414 wants to merge 8 commits into
Conversation
Independent reviewer terminal verdict: REQUEST_CHANGESReviewed candidate
Blocking finding: valid empty CTC hypotheses are excluded from WER/CER
Independent probe:
The existing Owner handoff
Independent validation on this exact head
Approval is blocked until the empty-hypothesis metric path is repaired and the replacement head is independently revalidated. |
|
Addressed on exact head Root cause and fix
Edge regression and real L3 evidence
Validation and public closure
The canonical PR body has been regenerated with this final-head evidence, exact public sequence, and Lane A #254 binding at |
Independent successor reviewer terminal verdict: APPROVEReviewed candidate
Prior blocker: CLOSED
Engineering review
Independent validation
Public and evidence closure
No actionable blocker remains. This is a skill-level reviewer opinion posted as a normal issue comment only; it does not change GitHub Review state or authorize merge/readiness changes. |
|
Successor independent review requested for exact head Sealed public evidence is |
Independent successor reviewer terminal verdict: APPROVEReviewed candidate
Engineering review
Independent validation
Public and evidence closure
No actionable blocker remains. This is a skill-level reviewer opinion posted as one normal issue comment only; it does not change GitHub Review state, labels, threads, readiness, or merge state. |
|
Successor evidence update for the prior exact-head review at #1335 (comment). The reviewed
The pinned Polish compatibility probe now agrees across modes at IDs Exact-head public closure is The canonical body now records |
Independent reviewer terminal verdict: APPROVEReviewed candidate
Engineering findings
Independent validation
Live closure
This is a skill-level reviewer opinion posted as one normal PR comment only. It does not submit GitHub Review state, edit the PR, resolve threads, mark ready, merge, or claim #1207 acceptance. |
|
Exact-head dependency closure update for This successor closes the incomplete selected-row identity validation defect reported by the independent reviewer on stacked Draft PR #1343: all selected rows are now fully validated before transcription/model inference as a nonnegative integer source index, a valid JSON-stable scalar semantic dataset ID, and a stable redacted audio key or basename. Source indices and normalized media identities must be unique; semantic dataset IDs are deliberately not uniqueness keys, so legitimate duplicate IDs remain valid when source and media identities differ. The 19 malformed/duplicate discriminators all recorded zero inference calls, while the valid duplicate-semantic-ID case processed both rows with preserved source order, redacted provenance, and requested/selected/processed/rejected/skipped accounting. Public exact-head evidence is The canonical body now binds the refined existing |
|
Independent review of exact head [P1] Honor
Please apply the seeded shuffle before selection for both streaming and non-streaming datasets while retaining each row's original source index, then add focused tests that distinguish All other reviewed exact-head gates passed: composite identity rejection before inference ( REQUEST_CHANGES |
|
Evidence reply to reviewer comment 5389680910 for exact head The blocking seeded-shuffle gap reported against Exact public-checkout discriminators:
Fresh sealed closure from the public fork ref and detached exact SHA:
Dependency-aware L0/L1/L2/Analyze evidence was reused without being relabeled fresh: execution provenance Lane A was refined in place rather than duplicated: draft gim-home/ModelKitArtifacts#254 at exact head The canonical PR body has been refreshed with these exact values. At this snapshot, all 9 exact-head checks are completed/successful and the sole review thread is resolved. The PR remains Draft with |
Independent reviewer terminal verdict: APPROVEReviewed candidate
Engineering findings
Independent validation
Live closure
No actionable blocker remains. This is a skill-level reviewer opinion posted as one normal PR comment only; it does not submit GitHub Review state, edit metadata or threads, mark ready, merge, or authorize a non-Draft transition. |
3402b7b to
02e51c2
Compare
|
APPROVE Independent reviewer verdict for exact published D5
This is a normal skill-level reviewer comment only. It does not authorize or perform a GitHub Review, body/thread/metadata mutation, ready transition, merge, or branch mutation. |
02e51c2 to
4b52813
Compare
|
APPROVE
No blocking finding remains. Hardware/model-scale measurements were not re-executed in this public terminal review; approval relies on the independently sealed artifacts under their explicitly preserved original commit provenance, with fresh exact-D6 evaluator, focused, static, and complete non-hardware partition validation above. |
Summary
This follow-up to merged PR #1177 adds shared, metadata-gated CTC automatic speech recognition evaluation for
facebook/mms-1b-all, fixes generic HF--no-optimizestage control, and retains verified CPU FP32/FP16 recipes. The contribution ships at Effort L2 / Outcome L2 and reaches the committed Goal L3 with CPU builds, performance, PyTorch parity, and bounded two-row FLEURS functional smokes. The final successor appliesDatasetConfig.shuffleandseedbefore the sample cap with deterministic bounded-buffer semantics, selects identical original source indices in streaming and non-streaming modes for a fixed seed, preservesshuffle=falsefirst-N Polish behavior, and retains the complete source-index/dataset-ID/redacted-media composite identity boundary before inference. The exact D6 candidate is4b528133fef8cbeed4e9a3aba9be42d22a3c0221on current main and merge basee28b128f5c2f69ecb2d73b63d2aea0a5ee8bddd0; it is a byte-equivalent replay of the accepted D5 series.Model metadata
What the model does
facebook/mms-1b-allis a 1 billion-parameter multilingual automatic speech recognition checkpoint. It accepts mono speech waveforms sampled at 16 kHz and emits frame-level CTC token logits that the checkpoint processor greedily decodes into text; published language adapters cover 1,162 languages.3d33597edbdaaba14a8e858e2c8caa76e3cec0cd, plus current-main WinML export metadata.verified.Primary user stories
AutoProcessor, argmax over logits, andprocessor.decode. Confidence:verified.tokenizer.set_target_langandmodel.load_adapteracross 1,162 supported languages. Confidence:verified.Supported tasks
automatic-speech-recognitionacross the checkpoint, Transformers, and WinML surfaces. Evidence: the checkpoint pipeline tag,Wav2Vec2ForCTCarchitecture, metadata-derived WinML registration forwav2vec2, and recipe-freewinml inspect/build resolution. Confidence:verified.Model architecture
Wav2Vec2ForCTCsource and meta-device module tree, and current-main export hierarchy metadata (verified).Validation and support evidence
1. Baseline
microsoft/winml-climain at0876e5ae1c98a169a6137e092e0d7b30bf9cee33, WinML0.3.0. L0/L2 values retain execution provenancecac2526b620df08fdbcc26ef60fe590479e55a0e; L1 Perf and Analyze retain C5 execution provenanceac8acd690b48828eb7dfdd38c4aa5de9e3bcb14dwith artifacts fromcac2526b620df08fdbcc26ef60fe590479e55a0e; L3 retains artifact executioncac2526b620df08fdbcc26ef60fe590479e55a0eand evaluator execution3440c5401d0111f06456b00662469daad2a6f0f6. Shipment identity is current main and merge basee28b128f5c2f69ecb2d73b63d2aea0a5ee8bddd0, exact D6 head4b528133fef8cbeed4e9a3aba9be42d22a3c0221, parent34bc2a0e6181e8c9f0cc27fddf9ebfaf2d1dbfe7, and treeb9b8745bb942e2f359cf583a642c3d8cc0112282; current main is its ancestor.b91152fb07a8a769a4756e94006bf83ede7073c9on base8079522ee692645569508d11abd4a0ea45a40987) established recipe-only CPU FP32/FP16 support and explicitly did not establish WER or benchmark accuracy. This is a new follow-up branch and PR; recipe(mms-1b-all): add verified CPU fp32/fp16 configs #1177 is not reopened or modified.audio-classification,audio-frame-classification,audio-xvector,automatic-speech-recognition, andfeature-extractionresolved. Verdict:WINML-ONLY.AutoModelForCTC/automatic-speech-recognitionand producedinput_valuesfloat32[1,16000]->logitsfloat32[1,49,154]. It also exposed the generic defect that Optimize ran despite explicit--no-optimizebecause the HF pipeline droppedskip_optimize.Task 'automatic-speech-recognition' is not supported.This set the Goal floor at L1 and established the shared CTC evaluator gap.Wav2Vec2ForCTC; the shipped recipes use metadata-derivedAutoModelForCTCand add the bounded Eval contract.2. Goal
L2, Goal ceilingL3, OutcomeL2; no ceiling downgrade or re-issued charter.CPUExecutionProvider / cpu, required precisions FP32 and FP16.3. Outcome
PASS, technical stateCLOSED, and dependency public acquisition stateCLOSED; all 14 fresh D6 public rows are closed and the independently rewalked 49-file seal matches byte-for-byte. Highest Goal verdict:L3 PASS.full; deferred tuples: none.examples/recipes/facebook_mms-1b-all/cpu/cpu/automatic-speech-recognition_fp32_config.jsonandexamples/recipes/facebook_mms-1b-all/cpu/cpu/automatic-speech-recognition_fp16_config.json.src/winml/modelkit/eval/ctc_asr_evaluator.py, evaluator registration/defaults/schema and generic wrapper registration, plus HF build stage-control forwarding insrc/winml/modelkit/commands/build.py.ac8acd690b48828eb7dfdd38c4aa5de9e3bcb14d; optim retains3440c5401d0111f06456b00662469daad2a6f0f6. Earlier full Eval and five-partition results remain retained evidence under their original provenance and are not relabeled as fresh D6 execution.wav2vec2-027now recordsDatasetConfig.shuffle/seedbefore bounded selection, deterministic streaming/non-streaming original-index equivalence, bounded finite-buffer consumption, pre-shuffle original-source provenance, and preserved composite source-index/dataset-ID/redacted-media identity. No duplicate finding was created.b7efe2b77c59eafbd13c48c4cb9beb9336df8305, labelmodel-scale-by-skill, checkGitOps/AdvancedSecurity: COMPLETED/SUCCESSon that exact head. Refined_meta-112binds the paired tester/reviewer contracts to the same shuffle-before-cap, fixed-seed mode-equivalence, bounded-consumption, original-provenance, and preserved composite-identity requirements.wav2vec2-027and_meta-112remain the durable binding for shuffle-before-cap, deterministic streaming/non-streaming original-index equivalence, bounded buffering, and composite row/media provenance; D6 refreshes candidate, VED-compatibility, public-acquisition, and seal provenance only, with no duplicate finding or new verdict shape.Final closure rows
[7, 4, 0, 2]in both modes; seed 456 selected[4, 6, 3, 2]; 1,200-row selection matched at[15, 682, 592, 53]; infinite-style source yielded 1,007 rows under ceiling 1,100DatasetValidationErrorwith 0 transcription calls and 0 model callsshuffle=false[1525, 1657], original indices[0, 1], and basenames10018492969996036091.wav,10288018704489549018.wavvision_encoder_decoder.py; no owned-path overlap or CTC reach; 15 targeted VED tests and the models partition (1,538 passed, 6 skipped, 2 xfailed) passedcac2526b; L1/Analyze at C5ac8acd69with artifacts atcac2526b; L3 atcac2526b+ evaluator3440c540; unaffected quality at C5/3440c540; none relabeled as fresh D6 executionaudioinitialization verified; all 14 fresh rows closed; independent 49-file rewalk matchedgit show --pretty=format: --no-ext-diff --binary <commit>togit patch-id --stablepipeline. Ordered IDs are2e0f73d16f95a28fce9ea36da84c2cc7fa68850e,e84028b1e13c416b39a9133ff930159e79689560,902bae581a224965e269eb3546a0bb9ad565049a,6e1039c7d747d5c39658d3601864a9ddf0312098,1661be1ec70cd0b1232133b14b912a99ee84e4db,6dfe21004adfba20ec0fe86966160008bd57ea1e,a09c8752c26cfb1609c68ef107ad003c0c0b3a51, and1fb43c8e1bb54e3fceca1c490d27fa9d9b49fa59. They match the accepted D5 and exact D6 series in order. Earlier serialized patch-ID fields had no generating command and are noncanonical historical evidence; they are not used for acceptance.4. Per-EP/device/precision results and Functional smoke Eval
FP16 is slower on this CPU; no speedup is claimed. Both runs used 20 iterations and 3 warmups with the same named float32
[1,16000]input.Functional smoke Eval:
PASSwith original public execution provenance at accepted head3402b7b4b1990a0932551a3d69d3838df0e70bda, carried to exact D64b528133fef8cbeed4e9a3aba9be42d22a3c0221by the byte-equivalent accepted series and 10/10 final-blob equality; it was not freshly re-executed on D6. The FP32 CPU smoke used a fresh exact-head MMS FP32 build from that accepted public checkout andgoogle/fleursrevision70bb2e84b976b7e960aa89f1c648e09c59f894dd, configen_us, validation split. Deterministic no-shuffle source-order selection processed source IDs[1548, 1620]at indices[0, 1], exactly 2 rows with 0 skipped and 0 rejected; emitted redacted row provenance, schema, label semantics, prediction semantics, and accounting were verified. Fan-out was capped at 2 dataset rows, batch size 1, at most 64 fixed waveform windows per utterance, no candidate-label/prompt expansion, no beams, and one full CTC decode across accepted windows. The processor-selected adapter language remainedeng, independently of the FLEURS localeen_us. Raw metrics were WER0.8541666666666666and CER0.3684210526315789.This is functional-smoke-only accounting: it proves real-audio decode, complete-utterance preprocessing, ONNX inference, CTC decoding, and metric emission. It is not representative accuracy or benchmark quality, and it does not claim accuracy for FP16 or other EPs. The former blocker was the absent ASR task path; the shared evaluator now supplies SoundFile bytes/path decoding, metadata-derived resampling/language/vocabulary handling, static/dynamic waveform processing, processor-owned CTC decode, WER/CER, and fail-closed row accounting.
5. Delta
Recipe delta from frozen auto-config
/loader/model_classWav2Vec2ForCTCAutoModelForCTC/evalnullgoogle/fleurs@70bb2e84b976b7e960aa89f1c648e09c59f894dd,en_us, validation, two rows, deterministic ordering,audio/transcription/loader/model_classWav2Vec2ForCTCAutoModelForCTC/quantnullmode=fp16,fp16_keep_io_types=true/evalnullgoogle/fleurs@70bb2e84b976b7e960aa89f1c648e09c59f894dd,en_us, validation, two rows, deterministic ordering,audio/transcriptionThe delta is reducible to metadata-derived CTC behavior plus the checkpoint's bounded dataset/precision recipe intent. Recipe-free inspect acceptance passed with
wav2vec2 -> automatic-speech-recognition -> AutoModelForCTC -> Wav2Vec2OnnxConfig -> WinMLModelForGenericTask. The production recipe README remains untouched. FP16 requires the explicit public--precision fp16build contract because CPU auto-precision otherwise overrides the recipe's FP16 mode.Shared CTC evaluator capability
winml evalfor metadata-resolved CTC ASR failed before dataset loading becauseautomatic-speech-recognitionhad no evaluator path.WinMLCTCASREvaluatorand its decode/resample/language/error-rate helpers implement SoundFile bytes/path decode, mono conversion, processor-rate resampling, deterministic static/dynamic utterance handling, vocabulary checks, processor-owned CTC decode, and fail-closed WER/CER accounting; registry/schema/default-dataset/generic-wrapper mappings route only metadata-resolved CTC models.en_usis not guessed to equaleng.Generic CTC processor and language compatibility repair
Wav2Vec2ProcessorWithLMand failed before greedy CTC evaluation when optionalpyctcdecodewas absent; separately, ordinaryWav2Vec2CTCTokenizer(target_lang=None)was rejected as though it were a misconfigured MMS adapter checkpoint.target_langtruthiness instead of positive adapter metadata to distinguish ordinary Wav2Vec2 from MMS-style language adapters._load_ctc_processorpreserves any successfully loaded processor and falls back to a composedAutoFeatureExtractor+AutoTokenizerWav2Vec2Processoronly for the exact missing-pyctcdecodeWav2Vec2ProcessorWithLMimport failure._configure_processor_languageaccepts null language for non-adapter checkpoints, while positive adapter metadata requires a valid active or configured language and fails closed otherwise.target_lang=None; MMS preserves activeeng, validates requested adapters, and rejects invalid or missing language. Vocabulary-width validation, 64-window fan-out, blank-boundary insertion, and empty-hypothesis WER/CER accounting are unchanged.Wav2Vec2Processor, 16 kHz, vocab 40, blank 0, decode-a, and no LM decoder; the pinned MMS probe preservedengand vocab 154 and failed closed for invalid/missing adapters. Thirteen focused compatibility/empty-hypothesis cases, all 50 CTC tests, 686 eval tests, 716 loader/build tests, and the five CI partitions passed.Generic seeded bounded selection, composite identity, provenance, and accounting repair
WinMLCTCASREvaluator.prepare_data()ignoredDatasetConfig.shuffleandseed; withshuffle=true,seed=123, and a bounded sample cap, both modes silently evaluated first-N rows and could emit valid WER/CER for the wrong corpus.shuffle=trueapplies the configured seed beforetake(samples)through the same finite buffered-shuffle semantics in streaming and non-streaming modes;shuffle=falseretains native first-N order. The selected original index, valid scalar dataset ID, and stable redacted media key/basename continue through output and are validated as a composite identity before transcription/model inference.shuffle=falsePolish selection remains IDs[1525, 1657], indices[0, 1], with the same two redacted basenames; accounting, WER/CER, MMSeng, ordinary Wav2Vec2 optional-LM fallback and nulltarget_lang, and successful empty-hypothesis deletion scoring are preserved.[7, 4, 0, 2]in both modes and seed 456 selected[4, 6, 3, 2]; both modes selected[15, 682, 592, 53]across 1,200 rows; the infinite-style source yielded 1,007 rows under a 1,100 ceiling. The focused selection/identity matrix passed 19/19, full CTC passed 50/50, full Eval passed 686/686, the five CI partitions passed 8,543 tests, and current-main/D5 full-format parity remainedPASS_NO_REGRESSIONwith the same 95 inherited files.Empty CTC hypothesis accounting repair
""and was rejected, removing its deletion errors from corpus WER/CER.WinMLCTCASREvaluator.computeused a post-decode truthiness guard even though actual decode/validation failures were already represented separately by_RejectedSampleError; the guard therefore conflated output quality with sample failure.skipped_samplesis explicitly0.["hello world", "good day"]with predictions["", "good day"]produce WER0.5, CER11/19(0.5789473684210527), 2 processed, 0 skipped, and 0 rejected. Two empty hypotheses produce WER/CER1.0with 2 processed and 0 skipped/rejected. Focused edge tests: 4 passed; full CTC evaluator: 25 passed; full eval unit suite: 661 passed.Generic
--no-optimizebug fix--no-optimizestill ran graph optimization._build_hf_pipelinedropped the CLIskip_optimizestate before the optimize stage.src/winml/modelkit/commands/build.py::_build_hf_pipelineforwards explicit stage-control state while preserving independent quantize and compile behavior.--no-optimizetests passed, and a real tiny Wav2Vec2 HF probe completed with Optimize0.0 swhile export remained enabled.D6 head
4b528133fef8cbeed4e9a3aba9be42d22a3c0221is the complete eight-commit accepted D5 series replayed byte-equivalently onto current maine28b128f5c2f69ecb2d73b63d2aea0a5ee8bddd0; its final commit changes onlyctc_asr_evaluator.pyand its unit test relative to parent34bc2a0e6181e8c9f0cc27fddf9ebfaf2d1dbfe7. The r6 proof is 8/8 equal commits, 8/8 equal ordered canonical stable patch IDs, 10/10 equal changed path statuses, and 10/10 equal final blobs. The sole main increment changessrc/winml/modelkit/models/hf/vision_encoder_decoder.py, has no owned-path overlap or CTC reach, and is closed by fresh exact-D6 VED, models-partition, CTC/build, Ruff, mypy, license, and format-parity evidence. L0/L1/L2/L3/Analyze and unaffected quality rows retain their recorded original execution SHAs and are not relabeled as fresh D6 execution. The final GitHub snapshot is 9/9 completed-success checks; zero review threads are unresolved; all 14 fresh public rows and the independently rewalked 49-file seal are closed.CodeQL superclass-initialization repair
WinMLCTCASREvaluator(config, model)left the inheritedpipeattribute absent, and CodeQL reportedpy/missing-call-to-initat the evaluator constructor.model,config, anddatainstead of invokingWinMLEvaluator.__init__; that omission skipped the base-ownedpipeinitialization.WinMLCTCASREvaluator.__init__initializes processor-specific state and callssuper().__init__(config, model), whileprepare_pipeline()returnsNoneto preserve the evaluator's intentional direct-model CTC path.6. Analyze summary - component level and op level
ANALYZE-PARTIAL-SUCCESS: Analyze returned accepted exit code1in 70.28099999995902 s because six requested EP/device rows had no shipped rule data; all 12 rows were present and parseable, with no errors or warnings. This is static rule analysis, not runtime execution.Component-level summary
Mapping gaps: 686 fusion-generated nodes lack retained component scope and remain unmapped rather than being assigned by adjacency. Final dropout is disabled in evaluation mode and has no ONNX node.
Op-level summary
Rule-less rows: CUDA/GPU, MIGraphX/GPU, Tensorrt/GPU, DML/GPU, CPU/CPU, and VitisAI/NPU. Six of 12 records had rule-declared support; records with errors: 0; records with warnings: 0.
7. Reproduce commands
The following is the exact public acquisition, environment, data preparation, quality, FP32/FP16 build, perf, L2 compare, L3 functional-smoke, rules acquisition, and Analyze sequence. It requires a new or empty
$Outand first verifies that the public branch advertises the exact candidate SHA.