Fix dub assembly failing under torchaudio 2.9 when TorchCodec is absent - #2379
tokutei58301-boop wants to merge 1 commit into
Conversation
torchaudio 2.9 routes torchaudio.load() through TorchCodec, which is not in uv.lock — so on every uv sync source install, dub assembly raises ImportError and every generate fails with dub_speech_missing even though the TTS segments rendered fine. The save side of this backend shift is already guarded in services/audio_io.py; this adds the matching fallback for the load side, reading the rendered PCM WAV segments via soundfile. Fixes debpalash#2378 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
|
[Medium risk] Adds fallback audio loading when a dependency is missing. The PR is not ready to merge because supported dub-generation paths still fail when TorchCodec is absent.
|
| wav, loaded_sr = torchaudio.load(entry[2]) | ||
| try: | ||
| wav, loaded_sr = torchaudio.load(entry[2]) | ||
| except ImportError: |
There was a problem hiding this comment.
Other segment reads still fail When torchaudio 2.9 runs without TorchCodec,
smart_fit and stretch_video reach an unguarded torchaudio.load() in _entry_num_samples() before this fallback, so generation still aborts. Cached-segment reuse and remote-segment import also load audio before this fallback; use the same soundfile fallback for those reads.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: debpalash/VoiceStudio/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughThe dub entry WAV loader now uses soundfile when ChangesDub audio loading
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~8 minutes Change: Bug fix · Severity of issue fixed: Medium Suggested reviewers: Merge Risk: ⚪ Minimal · up to The fallback supports dubbing assembly when 🚥 Pre-merge checks | ✅ 7 | ❌ 2❌ Failed checks (2 warnings)
✅ Passed checks (7 passed)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Consolidate community engine, workflow, dictation, and setup fixes with the Electron composer, sidebar, and voice UI. Fix review findings in engine residency, remote exports, backup cleanup, bounded compressed-audio decoding, and reference preprocessing. Preserve contributor credits in CHANGELOG.md and leave the app version unchanged. Supersedes #2325, #2338, #2368, #2377, #2379, #2380, #2383, #2384, #2387, #2390, #2391, #2392, #2393, #2395, #2400, #2401, #2402, #2409, #2410, and #2412.
|
Implemented and merged through #2419. Contributor credit is preserved in CHANGELOG.md. |
Fixes #2378
Problem
With the lockfile-resolved torchaudio 2.9.1,
POST /dub/generate/{job_id}renders every TTS segment fine but assembly always fails withdub_speech_missingon any freshuv syncsource install:torchaudio.load()now routes through TorchCodec (torchaudio._torchcodec.load_with_torchcodec), andtorchcodecis not inuv.lock, so the call raisesImportError.Full reproduction, environment, and the log trace are in #2378.
Root cause
The repo already handles the save side of the torchaudio 2.9 backend shift —
services/audio_io.pydocuments the sox → soundfile → TorchCodec migration and guardstorchaudio.save(that's why segment rendering works). The load side in_load_entry_wav()was not covered.Change
backend/api/routers/dub_generate.py—_load_entry_wav(): wrap thetorchaudio.load()call intry/except ImportErrorand fall back tosoundfile.read(dtype="float32", always_2d=True). The tensors loaded here are the dub's own renderedseg_*.wavfiles (plain PCM WAV written by the sibling soundfile save path), so the soundfile read is exactly equivalent; the existingtorchaudio.functional.resamplepath is untouched.soundfileis already a locked direct dependency (used byaudio_io.py), so no dependency changes.Verification
Same machine (Windows 11, RTX 5060 Ti), same job re-run end-to-end after the patch:
POST /dub/generate/{job_id}→data: {"type": "done", "segments_processed": 1, "language_code": "vi", "tracks": ["vi"], "sync_scores": [0.718], "fit_status": [{"status": "fits"}], ...}GET /dub/download/{job_id}→ two-audio-track MP4 (original +vidub)POST /transcribeon it confirms speech present and timed correctlyNotes for maintainers
#2378 also lists ~10 other
torchaudio.loadcall sites inbackend/with the same latent exposure (watermark.py,longform_render.py,persona_bundle.py,tts_backend.py, engine adapters, and 4 more in this file). This PR deliberately keeps the minimal fix to the path users hit today; if you'd prefer the durable variant instead — a sharedservices/audio_io.load_audio()helper with this fallback, migrated across all call sites, plus a smoke test that loads a rendered segment so a lockfile bump fails CI rather than users — I'm happy to rework this PR in that direction._load_entry_wav()now falls back tosoundfilewhentorchaudio.load()raisesImportError, while preserving channel-first tensor conversion and the existing resampling path. This lets dub assembly read rendered PCM WAV segments when TorchCodec is unavailable. Othertorchaudio.load()call sites remain unchanged, so the same failure may affect them.