Skip to content

Fix dub assembly failing under torchaudio 2.9 when TorchCodec is absent - #2379

Closed
tokutei58301-boop wants to merge 1 commit into
debpalash:mainfrom
tokutei58301-boop:fix/torchaudio-load-torchcodec-fallback
Closed

tokutei58301-boop wants to merge 1 commit into
debpalash:mainfrom
tokutei58301-boop:fix/torchaudio-load-torchcodec-fallback

Conversation

@tokutei58301-boop

@tokutei58301-boop tokutei58301-boop commented Sep 28, 2026 •

Copy link
Copy Markdown

Fixes #2378

Problem

With the lockfile-resolved torchaudio 2.9.1, POST /dub/generate/{job_id} renders every TTS segment fine but assembly always fails with dub_speech_missing on any fresh uv sync source install: torchaudio.load() now routes through TorchCodec (torchaudio._torchcodec.load_with_torchcodec), and torchcodec is not in uv.lock, so the call raises ImportError.

Full reproduction, environment, and the log trace are in #2378.

Root cause

The repo already handles the save side of the torchaudio 2.9 backend shift — services/audio_io.py documents the sox → soundfile → TorchCodec migration and guards torchaudio.save (that's why segment rendering works). The load side in _load_entry_wav() was not covered.

Change

backend/api/routers/dub_generate.py — _load_entry_wav(): wrap the torchaudio.load() call in try/except ImportError and fall back to soundfile.read(dtype="float32", always_2d=True). The tensors loaded here are the dub's own rendered seg_*.wav files (plain PCM WAV written by the sibling soundfile save path), so the soundfile read is exactly equivalent; the existing torchaudio.functional.resample path is untouched.

try:
    wav, loaded_sr = torchaudio.load(entry[2])
except ImportError:
    # torchaudio 2.9 routes load() through TorchCodec; when that wheel
    # is missing/broken, fall back to soundfile (segments here are
    # always plain PCM WAV written by the sibling soundfile save path).
    import soundfile as sf

    data, loaded_sr = sf.read(entry[2], dtype="float32", always_2d=True)
    wav = torch.from_numpy(data.T)

soundfile is already a locked direct dependency (used by audio_io.py), so no dependency changes.

Verification

Same machine (Windows 11, RTX 5060 Ti), same job re-run end-to-end after the patch:

  • POST /dub/generate/{job_id} → data: {"type": "done", "segments_processed": 1, "language_code": "vi", "tracks": ["vi"], "sync_scores": [0.718], "fit_status": [{"status": "fits"}], ...}
  • GET /dub/download/{job_id} → two-audio-track MP4 (original + vi dub)
  • Round-trip check: extracted the dub track, POST /transcribe on it confirms speech present and timed correctly

Notes for maintainers

#2378 also lists ~10 other torchaudio.load call sites in backend/ with the same latent exposure (watermark.py, longform_render.py, persona_bundle.py, tts_backend.py, engine adapters, and 4 more in this file). This PR deliberately keeps the minimal fix to the path users hit today; if you'd prefer the durable variant instead — a shared services/audio_io.load_audio() helper with this fallback, migrated across all call sites, plus a smoke test that loads a rendered segment so a lockfile bump fails CI rather than users — I'm happy to rework this PR in that direction.

_load_entry_wav() now falls back to soundfile when torchaudio.load() raises ImportError, while preserving channel-first tensor conversion and the existing resampling path. This lets dub assembly read rendered PCM WAV segments when TorchCodec is unavailable. Other torchaudio.load() call sites remain unchanged, so the same failure may affect them.

torchaudio 2.9 routes torchaudio.load() through TorchCodec, which is not
in uv.lock — so on every uv sync source install, dub assembly raises
ImportError and every generate fails with dub_speech_missing even though
the TTS segments rendered fine. The save side of this backend shift is
already guarded in services/audio_io.py; this adds the matching fallback
for the load side, reading the rendered PCM WAV segments via soundfile.

Fixes debpalash#2378

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
@greptile-apps

greptile-apps Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Retrigger

[Medium risk] Adds fallback audio loading when a dependency is missing.

The PR is not ready to merge because supported dub-generation paths still fail when TorchCodec is absent.

Fix All in Claude CodeFindings

  1. P1 Other segment reads still fail ▶
Summary

Adds a soundfile fallback when dub assembly cannot load a rendered segment through torchaudio.

  • Other segment-read paths in the same endpoint remain unguarded, leaving supported generation modes broken without TorchCodec.

Reviews (1) · Last reviewed commit: "Fix dub assembly failing under torchaudi..."

wav, loaded_sr = torchaudio.load(entry[2])
try:
wav, loaded_sr = torchaudio.load(entry[2])
except ImportError:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Other segment reads still fail When torchaudio 2.9 runs without TorchCodec, smart_fit and stretch_video reach an unguarded torchaudio.load() in _entry_num_samples() before this fallback, so generation still aborts. Cached-segment reuse and remote-segment import also load audio before this fallback; use the same soundfile fallback for those reads.

Fix in Claude Code

@coderabbitai

coderabbitai Bot commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: debpalash/VoiceStudio/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: a24ac537-4cfd-498a-b85d-c4f5f950dddb

📥 Commits

Reviewing files that changed from the base of the PR and between 08a1592 and e2bfcf2.

📒 Files selected for processing (1)
  • backend/api/routers/dub_generate.py

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

The dub entry WAV loader now uses soundfile when torchaudio.load raises ImportError. It converts the audio to a channel-first tensor. The existing sample-rate resampling remains unchanged.

Changes

Dub audio loading

Layer / File(s) Summary
Entry WAV load fallback
backend/api/routers/dub_generate.py
_load_entry_wav reads the WAV with soundfile when torchaudio.load raises ImportError. It converts the data to a channel-first tensor. Other load errors still propagate.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~8 minutes

Change: Bug fix · Severity of issue fixed: Medium

Suggested reviewers: debpalash

Merge Risk: ⚪ Minimal · up to e2bfc

The fallback supports dubbing assembly when torchaudio.load cannot load the rendered WAV; no material merge-blocking risk remains.

🚥 Pre-merge checks | ✅ 7 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Title check ⚠️ Warning The title accurately describes the fix and the body includes issue reference #2378, but it does not use the required Conventional Commit format with a scope, such as "fix(dub): …". Change the title to a scoped Conventional Commit title, for example: "fix(dub): fall back to soundfile when TorchCodec is unavailable".
Docstring Coverage ⚠️ Warning Docstring coverage is 33.33% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 3 functions across 1 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (7 passed)
Check name Status Explanation
Description check ✅ Passed The description provides the problem, root cause, implementation details, issue reference, and verification results. It omits the template's explicit Type and Checklist sections, but the required chan…
Linked Issues check ✅ Passed For #2378, _load_entry_wav() catches ImportError from torchaudio.load(), reads the PCM WAV with soundfile, and converts the result to a channel-first tensor. The existing resampling path remai…
Out of Scope Changes check ✅ Passed The reviewed diff changes only _load_entry_wav() in backend/api/routers/dub_generate.py. The fallback directly addresses the linked assembly failure and introduces no unrelated behavior or depende…
Cross-Platform Default Parity ✅ Passed No platform-divergent default is introduced. In the changed _load_entry_wav() path, every platform catches the same ImportError, calls the same `soundfile.read(..., dtype="float32", always_2d=True…
I18n Completeness (21 Locales) ✅ Passed The pull request changes only backend/api/routers/dub_generate.py. The diff contains no frontend files, no new or changed t('...') keys, and no changed frontend user-facing strings. The 21 locale …
Local-First Guarantee ✅ Passed PASS — The reviewed range changes only backend/api/routers/dub_generate.py. The added fallback performs local soundfile.read() and torch.from_numpy() operations. It adds no cloud call, account, …
Backward Compatibility ✅ Passed PASS: The PR changes only _load_entry_wav() in backend/api/routers/dub_generate.py. It adds a read fallback and does not change omnivoice_data, database schema, migrations, engine installation, …
  • Fix all pre-merge checks with AI

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

debpalash added a commit that referenced this pull request Sep 29, 2026
Consolidate community engine, workflow, dictation, and setup fixes with the Electron composer, sidebar, and voice UI.

Fix review findings in engine residency, remote exports, backup cleanup, bounded compressed-audio decoding, and reference preprocessing. Preserve contributor credits in CHANGELOG.md and leave the app version unchanged.

Supersedes #2325, #2338, #2368, #2377, #2379, #2380, #2383, #2384, #2387, #2390, #2391, #2392, #2393, #2395, #2400, #2401, #2402, #2409, #2410, and #2412.
@debpalash

Copy link
Copy Markdown
Owner

Implemented and merged through #2419. Contributor credit is preserved in CHANGELOG.md.

@debpalash debpalash closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Dub assembly fails with torchaudio 2.9: torchaudio.load routes through missing TorchCodec (dub_speech_missing)

3 participants