fix(asr): preserve merge token order at chunk seams (#825) - #830
Conversation
The batch long-form path merged overlapping windows with order-aware splicing, then globally re-sorted the merged stream by frame timestamp before building the transcript. That sort is destructive: - TDT emits several tokens per 80 ms frame, so many timestamps are equal. - The two overlapping windows' frame indices don't co-register across a seam, so a token that linearly follows another can carry a numerically smaller timestamp. convertTokensToText joins tokens in array order, so re-sorting by that coarse, non-monotonic key reorders same-frame subwords and interleaves subwords from the two windows — the reporter's German seam showed 'im Frühjahr' -> 'imüh Frjahr', 'Für die' -> 'die Für', 'Punkt' -> 'Pktun'. This is the token-order-inversion class #761 flagged as needing a fix in the merger itself. The pairwise merges already produce correct linear (text) order, so replace the global sort with enforceMonotonicTimestamps: clamp each backward-stepping timestamp up to the running max WITHOUT reordering. Text order is preserved; word timing and the seam-gap repair pass still see a non-decreasing sequence. collapseSeamWordDuplicates and the repair pass are unaffected (both rely on order + monotonic time, which the clamp keeps). Scope: fixes the merge-seam reordering. The first-chunk mel-context degradation (#594 family) the report also notes is separate. Adds token-level regression tests for the interleave/inversion cases.
Parakeet EOU Benchmark Results ✅Status: Benchmark passed Performance Metrics
Streaming Metrics
Test runtime: 0m50s • 07/31/2026, 10:57 PM EST RTFx = Real-Time Factor (higher is better) • Processing includes: Model inference, audio preprocessing, state management, and file I/O |
PocketTTS Smoke Test ✅
Runtime: 0m5s Note: PocketTTS uses CoreML MLState (macOS 15) KV cache + Mimi streaming state. CI VM lacks physical GPU — audio quality and performance may differ from Apple Silicon. |
Supertonic3 Smoke Test ✅
Runtime: 0m27s Note: CI VMs lack a physical Neural Engine; the ANE-bucketed VectorEstimator falls back to CPU here. This validates download + variant resolution + synthesis, not ANE residency/perf. |
VAD Benchmark ResultsPerformance Comparison
Dataset Details
✅: Average F1-Score above 70% |
Add the token-order-inversion failure mode, a "Token Order Is Authoritative" merge subsection explaining why timestamps are clamped monotonic instead of sorted, qualify the seam-garble known-limitation bullet (order-inversion subclass now fixed), and a changelog row.
Sortformer High-Latency Benchmark ResultsES2004a Performance (30.4s latency config)
Sortformer High-Latency • ES2004a • Runtime: 2m 47s • 2026-08-01T03:05:04.408Z |
Speaker Diarization Benchmark ResultsSpeaker Diarization PerformanceEvaluating "who spoke when" detection accuracy
Diarization Pipeline Timing BreakdownTime spent in each stage of speaker diarization
Speaker Diarization Research ComparisonResearch baselines typically achieve 18-30% DER on standard datasets
Note: RTFx shown above is from GitHub Actions runner. On Apple Silicon with ANE:
🎯 Speaker Diarization Test • AMI Corpus ES2004a • 1049.0s meeting audio • 41.0s diarization time • Test runtime: 2m 57s • 07/31/2026, 10:59 PM EST |
ASR Benchmark Results ✅Status: All benchmarks passed Parakeet v3 (multilingual)
Parakeet v2 (English-optimized)
Streaming (v3)
Streaming (v2)
Streaming tests use 5 files with 0.5s chunks to simulate real-time audio streaming 25 files per dataset • Test runtime: 9m38s • 07/31/2026, 11:10 PM EST RTFx = Real-Time Factor (higher is better) • Calculated as: Total audio duration ÷ Total processing time Expected RTFx Performance on Physical M1 Hardware:• M1 Mac: ~28x (clean), ~25x (other) Testing methodology follows HuggingFace Open ASR Leaderboard |
Offline VBx Pipeline ResultsSpeaker Diarization Performance (VBx Batch Mode)Optimal clustering with Hungarian algorithm for maximum accuracy
Offline VBx Pipeline Timing BreakdownTime spent in each stage of batch diarization
Speaker Diarization Research ComparisonOffline VBx achieves competitive accuracy with batch processing
Pipeline Details:
🎯 Offline VBx Test • AMI Corpus ES2004a • 1049.0s meeting audio • 100.3s processing • Test runtime: 1m 53s • 07/31/2026, 11:10 PM EST |
Fixes #825.
Root cause
The
AsrManagerbatch long-form path merges overlapping windows with order-aware, word-boundary-aware splicing (mergeChunks), which yields tokens in correct linear (text) order. It then globally re-sorted that stream by frame timestamp before building the transcript:convertTokensToTextjoins tokens in array order, so that sort determines the transcript. But frame timestamps are a bad sort key here:Re-sorting by that coarse, locally non-monotonic key reorders same-frame subwords and interleaves subwords from the two windows — exactly the reporter's German seam:
im Frühjahrimüh Frjahr(subwords interleaved)Für die Regaledie Für Regale(word order inverted)PunktPktun(same-frame subwords shuffled)This is the token-order-inversion class #761 flagged as needing "a fix in the merger itself."
Fix
The pairwise merges already produce correct order, so the global sort is not just unnecessary — it's the corruption. Replace it with
enforceMonotonicTimestamps, which clamps each backward-stepping timestamp up to the running maximum without reordering. Text order (the merger's output) is preserved; word timing and the opt-in seam-gap repair pass still see a non-decreasing timestamp sequence.collapseSeamWordDuplicatesandrepairSeamGapsare unaffected — both depend on token order + monotonic time, which the clamp keeps (and both previously ran on the scrambled stream).Scope
Targets the merge-seam reordering (the issue's title class). The first-chunk mel-context degradation the report also notes (all-lowercase, spoken-punctuation not converted) is the separate #594 family and is out of scope here.
Tests
Adds
ChunkSeamTokenOrderTests— token-level regression covering the cross-window subword interleave, the word-order inversion, same-frame ties, and the already-monotonic no-op, each asserting order is preserved while timestamps become non-decreasing. (Local XCTest is unavailable in my env; validated the clamp logic + golden values with a standalone harness, and the FluidAudio library target builds clean.)