diff --git a/CHANGELOG.md b/CHANGELOG.md index a092ccf22..d226bf7a5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -39,6 +39,10 @@ metadata and the backend fallback mirror it. ### Added +- Track model licence evidence and unresolved commercial-use reviews, with CI coverage for catalogue dependencies and backend repository defaults (#2587) + +- About credits supported speech models and conversions, with upstream terms and required Higgs Audio and Llama notices (#2587) + - MCP agents can design a voice from a text description and reuse it by `profile_id` (`describe_voice`, `design_voice`) (#2368) — thanks @thelselutopia! - Dictation vocabulary hint in Settings → Dictation shortcut: names and jargon that Faster Whisper, MLX Whisper and OpenAI-compatible engines should expect (#2395) — thanks @m061i6! - Cheaper Inference is available as an optional LLM provider (#2325) — thanks @aiapienthusiast! @@ -54,6 +58,17 @@ metadata and the backend fallback mirror it. - Pull requests ask their opener and every commit author and co-author to sign a one-time Contributor License Agreement by comment; the contributing guide explains how (#2556) ### Changed +- Keep existing recipes, watch folders and remote compute out of Pro benefit lists, and stop describing voice cloning as a paid unlock (#2587) +- Keep private licensing-service operations out of the public client protocol draft (#2587) +- Clarify in Pro, export and enterprise text that the application licence does not grant model or generated-output rights, in all 21 languages (#2587) + +- Remove unsupported automatic MIT licence claims from bundled demo metadata and generation scripts (#2587) + +- Correct historical OmniVoice commercial-licence claims and require the rights gate as well as technical smoke tests for engine acceptance (#2587) + +- Include the application licence notice and T3 Code MIT notice in desktop installers (#2587) + +- Remove the unused Remotion player integration and dependency while keeping the existing media playback providers (#2587) - Audio quality and Voice controls open as compact popovers from the Synthesize box on Clone and Voice Design (#2419) - The Synthesize button shows its keyboard shortcut as key chips inside the button (#2419) - The language menu stays within the window instead of clipping at the edges (#2419) @@ -87,6 +102,7 @@ metadata and the backend fallback mirror it. - Removed three hidden settings that nothing could set; the `OMNIVOICE_PRONUNCIATION` and `OMNIVOICE_TEXT_NORMALIZATION` switches remain (#2578) ### Docs +- Record the locked PyAV wheels' FFmpeg build flags, bundled codecs and unresolved redistribution terms (#2587) - Maintainer guide for repository settings that can't live in code; the licence notice scope and the contributing guide's list of network calls match the current app (#2556) - Contact addresses are now hi@voicestudio.sh (general and licensing), partner@voicestudio.sh (partnerships) and security@voicestudio.sh (security reports) (#2556) - Chinese README now matches the Electron installation and migration guide (#2377) — thanks @lg114! diff --git a/LICENSE-NOTICE.md b/LICENSE-NOTICE.md index 4215e66c7..4bf164dcf 100644 --- a/LICENSE-NOTICE.md +++ b/LICENSE-NOTICE.md @@ -55,9 +55,25 @@ weights as CC-BY-NC. Its `audio_tokenizer/LICENSE` contains separate Boson Higgs Audio 2 and Meta Llama community terms. A commercial license for VoiceStudio-owned code does not replace any of those terms. +The maintained About panel displays selected model credits and required literal +Higgs Audio and Llama attribution text. See [model credit sources](docs/model-credits.md) +for evidence and remaining gaps. These visible credits do not establish complete +notice compliance or permission for a particular use. + +The [model licence records](backend/config/model_licenses.json) distinguish +inspected non-commercial terms from unreviewed upstream metadata. A false +commercial-use flag includes unresolved review; it is not a claim that every +listed model forbids commercial use. No commercial clearance is asserted by +the initial inventory. + Third-party dependencies retain their own licenses. See `bun.lock`, `uv.lock`, and `native/desktop-bridge/Cargo.lock` for the resolved set. +The locked PyAV 15.1.0 wheels bundle FFmpeg and x264/x265 libraries. PyAV's source +licence alone does not describe those binaries' terms. The +[wheel audit](docs/licensing/pyav-15.1.0-audit.md) records their hashes, build flags, +upstream licence-label patch, and unresolved redistribution requirements. + ### Reference The full canonical text of the GNU Affero General Public License, Version 3 is diff --git a/backend/assets/samples/demo/dubbing/manifest.json b/backend/assets/samples/demo/dubbing/manifest.json index dc261ad89..852c49504 100644 --- a/backend/assets/samples/demo/dubbing/manifest.json +++ b/backend/assets/samples/demo/dubbing/manifest.json @@ -2,7 +2,8 @@ "version": "0.3.0", "rendered_by": "omnivoice engine + ffmpeg showwaves", "rendered_at": "2026-08-12T19:47:29Z", - "license": "MIT (synthetic, no third-party IP)", + "license": "NOASSERTION", + "license_review_required": true, "source": { "code": "en", "label": "English", diff --git a/backend/config/model_licenses.json b/backend/config/model_licenses.json new file mode 100644 index 000000000..00e7646f4 --- /dev/null +++ b/backend/config/model_licenses.json @@ -0,0 +1,931 @@ +{ + "schema_version": 1, + "checked_at": "2026-10-03", + "scope": "Configured model repositories and known dynamic asset families. commercial_use is reviewed clearance for this distribution, not a universal claim about the model; false includes unresolved terms. No runtime enforcement is implemented by this inventory.", + "models": [ + { + "id": "k2-fsa/OmniVoice", + "license": "CC-BY-NC (version unspecified); Higgs Audio 2; Meta Llama 3", + "commercial_use": false, + "review_status": "noncommercial", + "credit": "OmniVoice — Han Zhu and coauthors; tokenizer by Boson AI / Meta", + "source_url": "https://huggingface.co/k2-fsa/OmniVoice", + "evidence_url": "https://huggingface.co/k2-fsa/OmniVoice/blob/c5fdb5ccb189668d56333f77ba2629f4cd7535f4/README.md", + "revision": "c5fdb5ccb189668d56333f77ba2629f4cd7535f4", + "engines": [ + "omnivoice", + "omnivoice-subprocess" + ], + "configured_at": "backend/config/models.yaml", + "notes": "The pinned card distinguishes Apache-2.0 code from non-commercial weights. The audio tokenizer has separate Boson and Meta terms; see docs/model-credits.md." + }, + { + "id": "audio-cpp/audio.cpp-gguf", + "license": "other", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "audio-cpp/audio.cpp-gguf", + "source_url": "https://huggingface.co/audio-cpp/audio.cpp-gguf", + "evidence_url": "https://huggingface.co/audio-cpp/audio.cpp-gguf/blob/351dbab8d8534675ee29440bb402e348b09e55e2/README.md", + "revision": "351dbab8d8534675ee29440bb402e348b09e55e2", + "engines": [ + "audiocpp" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Systran/faster-whisper-large-v3", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-large-v3", + "source_url": "https://huggingface.co/Systran/faster-whisper-large-v3", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-large-v3/blob/edaa852ec7e145841d8ffdb056a99866b5f0a478/README.md", + "revision": "edaa852ec7e145841d8ffdb056a99866b5f0a478", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/whisper-large-v3-mlx", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/whisper-large-v3-mlx", + "source_url": "https://huggingface.co/mlx-community/whisper-large-v3-mlx", + "evidence_url": "https://huggingface.co/mlx-community/whisper-large-v3-mlx/blob/49e6aa286ad60c14352c404340ded53710378a11/README.md", + "revision": "49e6aa286ad60c14352c404340ded53710378a11", + "engines": [ + "mlx-whisper" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/whisper-large-v3-turbo", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/whisper-large-v3-turbo", + "source_url": "https://huggingface.co/mlx-community/whisper-large-v3-turbo", + "evidence_url": "https://huggingface.co/mlx-community/whisper-large-v3-turbo/blob/a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb/README.md", + "revision": "a4aaeec0636e6fef84abdcbe3544cb2bf7e9f6fb", + "engines": [ + "mlx-whisper" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "openai/whisper-large-v3", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "openai/whisper-large-v3", + "source_url": "https://huggingface.co/openai/whisper-large-v3", + "evidence_url": "https://huggingface.co/openai/whisper-large-v3/blob/06f233fe06e710322aca913c1bc4249a0d71fce1/README.md", + "revision": "06f233fe06e710322aca913c1bc4249a0d71fce1", + "engines": [ + "pytorch-whisper" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/whisper-tiny-mlx", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/whisper-tiny-mlx", + "source_url": "https://huggingface.co/mlx-community/whisper-tiny-mlx", + "evidence_url": "https://huggingface.co/mlx-community/whisper-tiny-mlx/blob/6caf9c55601caafbe6508a8b0d216bdf4783c4e8/README.md", + "revision": "6caf9c55601caafbe6508a8b0d216bdf4783c4e8", + "engines": [ + "mlx-whisper" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "deepdml/faster-whisper-large-v3-turbo-ct2", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "deepdml/faster-whisper-large-v3-turbo-ct2", + "source_url": "https://huggingface.co/deepdml/faster-whisper-large-v3-turbo-ct2", + "evidence_url": "https://huggingface.co/deepdml/faster-whisper-large-v3-turbo-ct2/blob/4df90f75321148c3a29a9e2351b7ddf8f5b115a8/README.md", + "revision": "4df90f75321148c3a29a9e2351b7ddf8f5b115a8", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Systran/faster-distil-whisper-large-v3", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-distil-whisper-large-v3", + "source_url": "https://huggingface.co/Systran/faster-distil-whisper-large-v3", + "evidence_url": "https://huggingface.co/Systran/faster-distil-whisper-large-v3/blob/c3058b475261292e64a0412df1d2681c06260fab/README.md", + "revision": "c3058b475261292e64a0412df1d2681c06260fab", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Systran/faster-whisper-medium", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-medium", + "source_url": "https://huggingface.co/Systran/faster-whisper-medium", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-medium/blob/08e178d48790749d25932bbc082711ddcfdfbc4f/README.md", + "revision": "08e178d48790749d25932bbc082711ddcfdfbc4f", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Systran/faster-whisper-small", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-small", + "source_url": "https://huggingface.co/Systran/faster-whisper-small", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-small/blob/536b0662742c02347bc0e980a01041f333bce120/README.md", + "revision": "536b0662742c02347bc0e980a01041f333bce120", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Systran/faster-whisper-base", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-base", + "source_url": "https://huggingface.co/Systran/faster-whisper-base", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-base/blob/ebe41f70d5b6dfa9166e2c581c45c9c0cfc57b66/README.md", + "revision": "ebe41f70d5b6dfa9166e2c581c45c9c0cfc57b66", + "engines": [ + "faster-whisper", + "faster-whisper-isolated", + "whisperx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "nvidia/parakeet-tdt-0.6b-v3", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "nvidia/parakeet-tdt-0.6b-v3", + "source_url": "https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3", + "evidence_url": "https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3/blob/541d1f99c6b0c3cd0b11a95167540bb8edefd82b/README.md", + "revision": "541d1f99c6b0c3cd0b11a95167540bb8edefd82b", + "engines": [ + "nemo-parakeet" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "nvidia/parakeet-tdt-0.6b-v2", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "nvidia/parakeet-tdt-0.6b-v2", + "source_url": "https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2", + "evidence_url": "https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2/blob/ae9ad07059c7c739ffaf932226a8fe64ae2620b0/README.md", + "revision": "ae9ad07059c7c739ffaf932226a8fe64ae2620b0", + "engines": [ + "nemo-parakeet" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/parakeet-tdt-0.6b-v3", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/parakeet-tdt-0.6b-v3", + "source_url": "https://huggingface.co/mlx-community/parakeet-tdt-0.6b-v3", + "evidence_url": "https://huggingface.co/mlx-community/parakeet-tdt-0.6b-v3/blob/ed2b7e8c15f9aaa0b5772e2efb986255eaef7e15/README.md", + "revision": "ed2b7e8c15f9aaa0b5772e2efb986255eaef7e15", + "engines": [ + "parakeet-mlx" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "UsefulSensors/moonshine-base", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "UsefulSensors/moonshine-base", + "source_url": "https://huggingface.co/UsefulSensors/moonshine-base", + "evidence_url": "https://huggingface.co/UsefulSensors/moonshine-base/blob/7a73d8d55ac0ba2ef3ae761593f6784b51f96dcf/README.md", + "revision": "7a73d8d55ac0ba2ef3ae761593f6784b51f96dcf", + "engines": [ + "moonshine" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "UsefulSensors/moonshine-tiny", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "UsefulSensors/moonshine-tiny", + "source_url": "https://huggingface.co/UsefulSensors/moonshine-tiny", + "evidence_url": "https://huggingface.co/UsefulSensors/moonshine-tiny/blob/390624ed33d594443aa4aa221f5b9f283b545b5a/README.md", + "revision": "390624ed33d594443aa4aa221f5b9f283b545b5a", + "engines": [ + "moonshine" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8/blob/2bda32ec70b097a55adaa07d9a7173915b43cc78/README.md", + "revision": "2bda32ec70b097a55adaa07d9a7173915b43cc78", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review. The raw card could not be retrieved (401/404); metadata alone does not resolve the access gap." + }, + { + "id": "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v2-int8/blob/1ab9323565ddb038682214b292f588070a538ce2/README.md", + "revision": "1ab9323565ddb038682214b292f588070a538ce2", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-streaming-zipformer-bilingual-zh-en-2023-02-20", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-streaming-zipformer-bilingual-zh-en-2023-02-20", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-bilingual-zh-en-2023-02-20", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-bilingual-zh-en-2023-02-20/blob/98590b7ed6443e77b714204da2757d75e1a642f4/README.md", + "revision": "98590b7ed6443e77b714204da2757d75e1a642f4", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-streaming-paraformer-bilingual-zh-en", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-streaming-paraformer-bilingual-zh-en", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-paraformer-bilingual-zh-en", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-paraformer-bilingual-zh-en/blob/8e40c43232a1c5c66c82111efc5820d3accca11b/README.md", + "revision": "8e40c43232a1c5c66c82111efc5820d3accca11b", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-streaming-zipformer-en-20M-2023-02-17", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-streaming-zipformer-en-20M-2023-02-17", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-en-20M-2023-02-17", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-en-20M-2023-02-17/blob/d42f2d9f7ca24806fb667456a18a9f1b60f70d16/README.md", + "revision": "d42f2d9f7ca24806fb667456a18a9f1b60f70d16", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-streaming-zipformer-zh-14M-2023-02-23/blob/204ad334e2e683fd295359930cc16fc0432a23ac/README.md", + "revision": "204ad334e2e683fd295359930cc16fc0432a23ac", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "csukuangfj/sherpa-onnx-whisper-tiny", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "csukuangfj/sherpa-onnx-whisper-tiny", + "source_url": "https://huggingface.co/csukuangfj/sherpa-onnx-whisper-tiny", + "evidence_url": "https://huggingface.co/csukuangfj/sherpa-onnx-whisper-tiny/blob/65176e2deb88badc814a94058666cadccc29b61c/README.md", + "revision": "65176e2deb88badc814a94058666cadccc29b61c", + "engines": [ + "sherpa-onnx-asr" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "facebook/nllb-200-distilled-600M", + "license": "CC-BY-NC-4.0", + "commercial_use": false, + "review_status": "noncommercial", + "credit": "NLLB-200 — Meta AI", + "source_url": "https://huggingface.co/facebook/nllb-200-distilled-600M", + "evidence_url": "https://huggingface.co/facebook/nllb-200-distilled-600M/blob/f8d333a098d19b4fd9a8b18f94170487ad3f821d/README.md", + "revision": "f8d333a098d19b4fd9a8b18f94170487ad3f821d", + "engines": [], + "configured_at": "backend/config/models.yaml", + "notes": "Pinned model card identifies CC-BY-NC-4.0. This record does not grant output or training-data rights." + }, + { + "id": "pyannote/speaker-diarization-3.1", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "pyannote/speaker-diarization-3.1", + "source_url": "https://huggingface.co/pyannote/speaker-diarization-3.1", + "evidence_url": "https://huggingface.co/pyannote/speaker-diarization-3.1/blob/84fd25912480287da0247647c3d2b4853cb3ee5d/README.md", + "revision": "84fd25912480287da0247647c3d2b4853cb3ee5d", + "engines": [], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "pyannote/segmentation-3.0", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "pyannote/segmentation-3.0", + "source_url": "https://huggingface.co/pyannote/segmentation-3.0", + "evidence_url": "https://huggingface.co/pyannote/segmentation-3.0/blob/e66f3d3b9eb0873085418a7b813d3b369bf160bb/README.md", + "revision": "e66f3d3b9eb0873085418a7b813d3b369bf160bb", + "engines": [], + "configured_at": "backend/config/models.yaml dependencies of pyannote/speaker-diarization-3.1", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "pyannote/wespeaker-voxceleb-resnet34-LM", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "pyannote/wespeaker-voxceleb-resnet34-LM", + "source_url": "https://huggingface.co/pyannote/wespeaker-voxceleb-resnet34-LM", + "evidence_url": "https://huggingface.co/pyannote/wespeaker-voxceleb-resnet34-LM/blob/837717ddb9ff5507820346191109dc79c958d614/README.md", + "revision": "837717ddb9ff5507820346191109dc79c958d614", + "engines": [], + "configured_at": "backend/config/models.yaml dependencies of pyannote/speaker-diarization-3.1", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "OpenMOSS-Team/MOSS-TTS-Nano-100M", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "OpenMOSS-Team/MOSS-TTS-Nano-100M", + "source_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-Nano-100M", + "evidence_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-Nano-100M/blob/44502f80dbf9743528fa921cc544d662c685ebec/README.md", + "revision": "44502f80dbf9743528fa921cc544d662c685ebec", + "engines": [ + "moss-tts-nano" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "KittenML/kitten-tts-mini-0.8", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "KittenML/kitten-tts-mini-0.8", + "source_url": "https://huggingface.co/KittenML/kitten-tts-mini-0.8", + "evidence_url": "https://huggingface.co/KittenML/kitten-tts-mini-0.8/blob/c02725660cea441db4c383af69f1f26f5cd00947/README.md", + "revision": "c02725660cea441db4c383af69f1f26f5cd00947", + "engines": [ + "kittentts" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "openbmb/VoxCPM2", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "openbmb/VoxCPM2", + "source_url": "https://huggingface.co/openbmb/VoxCPM2", + "evidence_url": "https://huggingface.co/openbmb/VoxCPM2/blob/32279effe8c19989596f05d353d1447f51d9e915/README.md", + "revision": "32279effe8c19989596f05d353d1447f51d9e915", + "engines": [ + "voxcpm2" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "FunAudioLLM/Fun-CosyVoice3-0.5B-2512", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "FunAudioLLM/Fun-CosyVoice3-0.5B-2512", + "source_url": "https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512", + "evidence_url": "https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512/blob/29e01c4e8d000f4bcd70751be16fa94bf3d85a18/README.md", + "revision": "29e01c4e8d000f4bcd70751be16fa94bf3d85a18", + "engines": [ + "cosyvoice" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "lj1995/GPT-SoVITS", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "lj1995/GPT-SoVITS", + "source_url": "https://huggingface.co/lj1995/GPT-SoVITS", + "evidence_url": "https://huggingface.co/lj1995/GPT-SoVITS/blob/336b2ec4e8d4ac74740798dd40af44e74659ecaf/README.md", + "revision": "336b2ec4e8d4ac74740798dd40af44e74659ecaf", + "engines": [ + "gpt-sovits" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/Kokoro-82M-bf16", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/Kokoro-82M-bf16", + "source_url": "https://huggingface.co/mlx-community/Kokoro-82M-bf16", + "evidence_url": "https://huggingface.co/mlx-community/Kokoro-82M-bf16/blob/a71e4d38b236d968966a2002c4c895dbd12b1c3c/README.md", + "revision": "a71e4d38b236d968966a2002c4c895dbd12b1c3c", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/csm-1b-8bit", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/csm-1b-8bit", + "source_url": "https://huggingface.co/mlx-community/csm-1b-8bit", + "evidence_url": "https://huggingface.co/mlx-community/csm-1b-8bit/blob/fcf0cc857eade3615a60f30722cf5197d4f88406/README.md", + "revision": "fcf0cc857eade3615a60f30722cf5197d4f88406", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit", + "source_url": "https://huggingface.co/mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit", + "evidence_url": "https://huggingface.co/mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-4bit/blob/5c390979e4b93af5f2932f90742ca99c7dd04687/README.md", + "revision": "5c390979e4b93af5f2932f90742ca99c7dd04687", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/Dia-1.6B", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/Dia-1.6B", + "source_url": "https://huggingface.co/mlx-community/Dia-1.6B", + "evidence_url": "https://huggingface.co/mlx-community/Dia-1.6B/blob/de4fa8c178ca5cc4e9d884b55b03fcfaa0995162/README.md", + "revision": "de4fa8c178ca5cc4e9d884b55b03fcfaa0995162", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/Llama-OuteTTS-1.0-1B-4bit", + "license": "CC-BY-NC-SA-4.0 AND Llama-3.2", + "commercial_use": false, + "review_status": "noncommercial", + "credit": "OuteAI; MLX conversion by mlx-community; Meta. Built with Llama.", + "source_url": "https://huggingface.co/mlx-community/Llama-OuteTTS-1.0-1B-4bit", + "evidence_url": "https://huggingface.co/mlx-community/Llama-OuteTTS-1.0-1B-4bit/blob/3ac2cff406f7de16a3216c60d0108571a916acc0/README.md", + "revision": "3ac2cff406f7de16a3216c60d0108571a916acc0", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Original model separates Llama 3.2 components from OuteAI non-commercial additions. Conversion card links original; see docs/model-credits.md." + }, + { + "id": "mlx-community/Chatterbox-TTS-4bit", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/Chatterbox-TTS-4bit", + "source_url": "https://huggingface.co/mlx-community/Chatterbox-TTS-4bit", + "evidence_url": "https://huggingface.co/mlx-community/Chatterbox-TTS-4bit/blob/a3c8ded2d711d6395410d645b3a97c79fd563a13/README.md", + "revision": "a3c8ded2d711d6395410d645b3a97c79fd563a13", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "mlx-community/MeloTTS-English-v3-MLX", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "mlx-community/MeloTTS-English-v3-MLX", + "source_url": "https://huggingface.co/mlx-community/MeloTTS-English-v3-MLX", + "evidence_url": "https://huggingface.co/mlx-community/MeloTTS-English-v3-MLX/blob/837d15fd72bc35a15033234ce5ea242367ca1960/README.md", + "revision": "837d15fd72bc35a15033234ce5ea242367ca1960", + "engines": [ + "mlx-audio" + ], + "configured_at": "backend/config/models.yaml", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "IndexTeam/IndexTTS-2.5", + "license": "other", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "IndexTeam/IndexTTS-2.5", + "source_url": "https://huggingface.co/IndexTeam/IndexTTS-2.5", + "evidence_url": "https://huggingface.co/IndexTeam/IndexTTS-2.5/blob/c39ce5ba981572cb187443877ff559dfb246ce63/README.md", + "revision": "c39ce5ba981572cb187443877ff559dfb246ce63", + "engines": [ + "indextts2" + ], + "configured_at": "backend/services/sidecar_install.py:411", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "Serveurperso/OmniVoice-GGUF", + "license": "CC-BY-NC-4.0; Higgs Audio 2; Meta Llama 3", + "commercial_use": false, + "review_status": "noncommercial", + "credit": "OmniVoice — Han Zhu and coauthors; GGUF conversion by Serveurperso; tokenizer by Boson AI / Meta", + "source_url": "https://huggingface.co/Serveurperso/OmniVoice-GGUF", + "evidence_url": "https://huggingface.co/Serveurperso/OmniVoice-GGUF/blob/017094167b5c9ed565a5076ac9b3b93c5ecf5c73/README.md", + "revision": "017094167b5c9ed565a5076ac9b3b93c5ecf5c73", + "engines": [ + "omnivoice-gguf" + ], + "configured_at": "backend/engines/omnivoice_gguf/backend.py:76", + "notes": "Conversion card names CC-BY-NC-4.0; its codec description does not override the actual tokenizer terms. See docs/model-credits.md." + }, + { + "id": "Supertone/supertonic-3", + "license": "openrail", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Supertone/supertonic-3", + "source_url": "https://huggingface.co/Supertone/supertonic-3", + "evidence_url": "https://huggingface.co/Supertone/supertonic-3/blob/3cadd1ee6394adea1bd021217a0e650ede09a323/README.md", + "revision": "3cadd1ee6394adea1bd021217a0e650ede09a323", + "engines": [ + "supertonic3" + ], + "configured_at": "backend/engines/supertonic3/constants.py:31", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "kyutai/pocket-tts", + "license": "cc-by-4.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "kyutai/pocket-tts", + "source_url": "https://huggingface.co/kyutai/pocket-tts", + "evidence_url": "https://huggingface.co/kyutai/pocket-tts/blob/3e82814a68665eec246ff649b14c71331f955c06/README.md", + "revision": "3e82814a68665eec246ff649b14c71331f955c06", + "engines": [ + "pockettts" + ], + "configured_at": "backend/engines/pockettts/main.py:215; SDK selects language/layer variants", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review. The raw card could not be retrieved (401/404); metadata alone does not resolve the access gap." + }, + { + "id": "OpenMOSS-Team/MOSS-TTS-v1.5", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "OpenMOSS-Team/MOSS-TTS-v1.5", + "source_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-v1.5", + "evidence_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-v1.5/blob/cdd3b911b1585e3f2dbc7775ef10f9926f58850a/README.md", + "revision": "cdd3b911b1585e3f2dbc7775ef10f9926f58850a", + "engines": [ + "moss-tts-v15" + ], + "configured_at": "backend/engines/moss_tts_v15/main.py:55", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "rednote-hilab/dots.tts-soar", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "rednote-hilab/dots.tts-soar", + "source_url": "https://huggingface.co/rednote-hilab/dots.tts-soar", + "evidence_url": "https://huggingface.co/rednote-hilab/dots.tts-soar/blob/2f9b3e18d70d670d4c701da2dc55ded5755815ce/README.md", + "revision": "2f9b3e18d70d670d4c701da2dc55ded5755815ce", + "engines": [ + "dots-tts" + ], + "configured_at": "backend/engines/dots_tts/main.py:53", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "netease-youdao/Confucius4-TTS", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "netease-youdao/Confucius4-TTS", + "source_url": "https://huggingface.co/netease-youdao/Confucius4-TTS", + "evidence_url": "https://huggingface.co/netease-youdao/Confucius4-TTS/blob/c82f9a7b3706ee3172eb95273daa650a504dccab/README.md", + "revision": "c82f9a7b3706ee3172eb95273daa650a504dccab", + "engines": [ + "confucius4-tts" + ], + "configured_at": "docs/engines/confucius4-tts.md:56", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "facebook/w2v-bert-2.0", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "facebook/w2v-bert-2.0", + "source_url": "https://huggingface.co/facebook/w2v-bert-2.0", + "evidence_url": "https://huggingface.co/facebook/w2v-bert-2.0/blob/da985ba0987f70aaeb84a80f2851cfac8c697a7b/README.md", + "revision": "da985ba0987f70aaeb84a80f2851cfac8c697a7b", + "engines": [ + "confucius4-tts" + ], + "configured_at": "docs/engines/confucius4-tts.md:59", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "funasr/campplus", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "funasr/campplus", + "source_url": "https://huggingface.co/funasr/campplus", + "evidence_url": "https://huggingface.co/funasr/campplus/blob/e4b6ede7ce16997aff4ae69fbca1f0175e2afede/README.md", + "revision": "e4b6ede7ce16997aff4ae69fbca1f0175e2afede", + "engines": [ + "confucius4-tts" + ], + "configured_at": "docs/engines/confucius4-tts.md:60", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "nvidia/bigvgan_v2_22khz_80band_256x", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "nvidia/bigvgan_v2_22khz_80band_256x", + "source_url": "https://huggingface.co/nvidia/bigvgan_v2_22khz_80band_256x", + "evidence_url": "https://huggingface.co/nvidia/bigvgan_v2_22khz_80band_256x/blob/633ff708ed5b74903e86ff1298cf4a98e921c513/README.md", + "revision": "633ff708ed5b74903e86ff1298cf4a98e921c513", + "engines": [ + "confucius4-tts" + ], + "configured_at": "docs/engines/confucius4-tts.md:61", + "notes": "Licence identifier is upstream model-card metadata, not verified permission for this distribution. Component terms and attribution text require review." + }, + { + "id": "iic/SenseVoiceSmall", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "iic/SenseVoiceSmall", + "source_url": "https://modelscope.cn/models/iic/SenseVoiceSmall", + "evidence_url": "https://modelscope.cn/models/iic/SenseVoiceSmall", + "revision": null, + "engines": [ + "funasr" + ], + "configured_at": "backend/services/asr_backend.py:2259", + "notes": "Runtime uses the FunASR/ModelScope identifier. The matching Hugging Face API returned 401; no governing weight terms have been verified." + }, + { + "id": "Systran/faster-whisper-tiny", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-tiny", + "source_url": "https://huggingface.co/Systran/faster-whisper-tiny", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-tiny/blob/d90ca5fe260221311c53c58e660288d3deb8d356/README.md", + "revision": "d90ca5fe260221311c53c58e660288d3deb8d356", + "engines": [ + "faster-whisper" + ], + "configured_at": "backend/services/asr_backend.py _fw_repo", + "notes": "Runtime fallback outside models.yaml. Licence identifier is model-card metadata; component terms and attribution require review." + }, + { + "id": "Systran/faster-whisper-large-v2", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Systran/faster-whisper-large-v2", + "source_url": "https://huggingface.co/Systran/faster-whisper-large-v2", + "evidence_url": "https://huggingface.co/Systran/faster-whisper-large-v2/blob/f0fe81560cb8b68660e564f55dd99207059c092e/README.md", + "revision": "f0fe81560cb8b68660e564f55dd99207059c092e", + "engines": [ + "faster-whisper" + ], + "configured_at": "backend/services/asr_backend.py _fw_repo", + "notes": "Runtime fallback outside models.yaml. Licence identifier is model-card metadata; component terms and attribution require review." + }, + { + "id": "openai/whisper-large-v3-turbo", + "license": "mit", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "OpenAI — Whisper large-v3-turbo", + "source_url": "https://huggingface.co/openai/whisper-large-v3-turbo", + "evidence_url": "https://huggingface.co/openai/whisper-large-v3-turbo/blob/41f01f3fe87f28c78e2fbf8b568835947dd65ed9/README.md", + "revision": "41f01f3fe87f28c78e2fbf8b568835947dd65ed9", + "engines": [ + "pytorch-whisper" + ], + "configured_at": "backend/services/tts_backend.py OMNIVOICE_PYTORCH_ASR_MODEL fallback", + "notes": "Clone-reference ASR fallback outside models.yaml. Model-card licence metadata is recorded; component terms and required notices remain under review." + }, + { + "id": "OpenMOSS-Team/MOSS-TTS-Nano", + "license": "apache-2.0", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "OpenMOSS-Team/MOSS-TTS-Nano", + "source_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-Nano", + "evidence_url": "https://huggingface.co/OpenMOSS-Team/MOSS-TTS-Nano/blob/44502f80dbf9743528fa921cc544d662c685ebec/README.md", + "revision": "44502f80dbf9743528fa921cc544d662c685ebec", + "engines": [ + "moss-tts-nano" + ], + "configured_at": "backend/services/tts_backend.py OMNIVOICE_MOSS_TTS_MODEL", + "notes": "Licence label is upstream metadata only; governing component terms, author credit and required notices remain under review." + }, + { + "id": "eustlb/higgs-audio-v2-tokenizer", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "eustlb/higgs-audio-v2-tokenizer", + "source_url": "https://huggingface.co/eustlb/higgs-audio-v2-tokenizer", + "evidence_url": "https://huggingface.co/eustlb/higgs-audio-v2-tokenizer/blob/528e871c2a26c4f0f7773b9754e2e1acae20899d/README.md", + "revision": "528e871c2a26c4f0f7773b9754e2e1acae20899d", + "engines": [ + "omnivoice" + ], + "configured_at": "backend/services/hf_revisions.py CURATED_REVISIONS; backend/services/model_manager.py repair", + "notes": "Licence label is upstream metadata only; governing component terms, author credit and required notices remain under review." + } + ], + "dynamic_assets": [ + { + "id": "OmniVoice text/audio tokenizer components", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://huggingface.co/k2-fsa/OmniVoice", + "evidence_url": "https://huggingface.co/k2-fsa/OmniVoice", + "configured_at": "omnivoice/models/omnivoice.py:526-545; LICENSE-NOTICE.md", + "notes": "Default repo carries audio_tokenizer/LICENSE; notice already identifies separate Higgs Audio 2 and Meta Llama terms. Review all tokenizer assets at resolved model revision." + }, + { + "id": "audioseal_wm_16bits / audioseal_detector_16bits", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/facebookresearch/audioseal", + "evidence_url": "https://github.com/facebookresearch/audioseal", + "configured_at": "backend/services/watermark.py:196,213", + "notes": "SDK resolves checkpoint URLs; record generator and detector weight terms separately from SDK." + }, + { + "id": "htdemucs", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/facebookresearch/demucs", + "evidence_url": "https://github.com/facebookresearch/demucs", + "configured_at": "backend/services/dub_pipeline.py:1537; backend/api/routers/system.py:1364", + "notes": "Demucs resolves actual checkpoint; record bundled/runtime weight revision and terms." + }, + { + "id": "WhisperX language-specific alignment models and silero VAD", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/m-bain/whisperX", + "evidence_url": "https://github.com/m-bain/whisperX", + "configured_at": "backend/services/asr_backend.py:576,818", + "notes": "SDK language mapping selects models; enumerate resolved installed SDK mapping, not only ASR catalogue repos." + }, + { + "id": "fsmn-vad / cam++", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/modelscope/FunASR", + "evidence_url": "https://github.com/modelscope/FunASR", + "configured_at": "backend/services/asr_backend.py:2260,2263", + "notes": "FunASR aliases resolve independently of SenseVoice; inspect SDK mappings." + }, + { + "id": "Argos language-pair packages", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/argosopentech/argospm-index", + "evidence_url": "https://github.com/argosopentech/argospm-index", + "configured_at": "backend/services/translation_engines.py", + "notes": "Language package selection dynamic; capture package identity/version, URL, hashes, terms at install." + }, + { + "id": "User-provided Sherpa ONNX / RVC voice models / GPT-SoVITS server / custom MLX and env-overridden repos", + "license": "NOASSERTION", + "commercial_use": false, + "review_status": "unreviewed", + "credit": "Unresolved: record the author and required notice for each selected asset.", + "source_url": "https://github.com/debpalash/VoiceStudio/blob/main/docs/engine-acceptance.md", + "evidence_url": "https://github.com/debpalash/VoiceStudio/blob/main/docs/engine-acceptance.md", + "configured_at": "backend/services/tts_backend.py; backend/services/rvc.py; backend/services/asr_backend.py", + "notes": "Cannot enumerate every user asset from repository; represent unknown/unreviewed state explicitly, with user provenance inputs." + } + ] +} diff --git a/backend/config/models.yaml b/backend/config/models.yaml index 338d931dd..2bce29313 100644 --- a/backend/config/models.yaml +++ b/backend/config/models.yaml @@ -4,6 +4,8 @@ # The backend loads it at startup via `load_model_catalog()`. # # To add a model: append an entry with the fields below. +# Add its licence/evidence record to model_licenses.json in the same change; +# CI checks top-level entries, nested dependencies and source-code defaults. # To remove: delete the entry. The UI will stop showing it immediately. # # Fields: diff --git a/bun.lock b/bun.lock index 8e6162fad..edc850b58 100644 --- a/bun.lock +++ b/bun.lock @@ -64,7 +64,6 @@ "react-hot-toast": "^2.6.0", "react-i18next": "^17.0.14", "react-window": "^2.3.1", - "remotion": "4.0.524", "shadcn": "^4.21.0", "sonner": "^2.0.8", "tailwind-merge": "^3.6.0", @@ -2082,8 +2081,6 @@ "remark-stringify": ["remark-stringify@11.0.0", "", { "dependencies": { "@types/mdast": "^4.0.0", "mdast-util-to-markdown": "^2.0.0", "unified": "^11.0.0" } }, "sha512-1OSmLd3awB/t8qdoEOMazZkNsfVTeY4fTsgzcQFdXNq8ToTN4ZGwrMnlda4K6smTFKD+GRV6O48i6Z4iKgPPpw=="], - "remotion": ["remotion@4.0.524", "", { "peerDependencies": { "react": ">=16.8.0", "react-dom": ">=16.8.0" } }, "sha512-jtoQbO7+UD7/4gcl0Onjq/Q27DP3qjI9hRimJJGuvU9p6OekY+Oyn1wNjYxG+hGH4i6InDFneN3sazFLuSF4Og=="], - "require-directory": ["require-directory@2.1.1", "", {}, "sha512-fGxEI7+wsG9xrvdjsrlmL22OMTTiHRwAMroiEeMgq8gzoLC/PQr7RsRDSTLUg/bZAZtF+TVIkHc6/4RIKrui+Q=="], "require-from-string": ["require-from-string@2.0.2", "", {}, "sha512-Xf0nWe6RseziFMu+Ap9biiUbmplq6S9/p+7w7YXP/JBHhrUDDUhwa+vANyubuqfZWTveU//DYVGsDG7RKL/vEw=="], diff --git a/docs/adr/SPIKE-01-gguf-research.md b/docs/adr/SPIKE-01-gguf-research.md index f50e63498..2d78cde00 100644 --- a/docs/adr/SPIKE-01-gguf-research.md +++ b/docs/adr/SPIKE-01-gguf-research.md @@ -3,21 +3,41 @@ # Phase 4: Adaptive & Specialty Engines (spike-first) — Research +## Licence correction (2026-10-03) + +This correction supersedes the original commercial-compatibility conclusions below. +The [upstream OmniVoice card](https://huggingface.co/k2-fsa/OmniVoice#license) +distinguishes Apache-2.0 code from CC-BY-NC pretrained weights. The +[GGUF card](https://huggingface.co/Serveurperso/OmniVoice-GGUF#license) now lists +CC-BY-NC-4.0 for its converted weights. The +[singing card](https://huggingface.co/ModelsLab/omnivoice-singing) still advertises +Apache-2.0 and identifies OmniVoice as its base; that label does not resolve the +conflicting upstream terms. Commercial clearance for the singing weights remains +unverified. Tokenizer/component terms also require separate review; see +[LICENSE-NOTICE.md](../../LICENSE-NOTICE.md). + +These historical engineering decisions do not approve any model for Pro. Model +rights must be recorded separately from runtime-code rights, with unresolved or +non-commercial weights excluded from Pro unless the necessary rights are obtained. +The required acceptance and backend enforcement work is tracked in #2587; this +correction does not claim those controls are already implemented. + + **Researched:** 2026-05-18 **Domain:** TTS engine integration — quantized GGUF runtime + singing-voice variant, both descending from the same `k2-fsa/OmniVoice` lineage already shipping as VoiceStudio's default cloning engine -**Confidence:** HIGH for model identity, license, runtime requirements, and GO/NO-GO calls (model cards confirmed, lineage chain verified end-to-end). MEDIUM for performance/latency numbers (no public benchmarks). MEDIUM for the heuristic singing/spoken segmentation strategy (SING-03 — feasible from existing toolkit but unbenchmarked). +**Confidence:** HIGH for model identity and runtime requirements. Original licence and commercial GO/NO-GO conclusions are superseded by the correction above. MEDIUM for performance/latency numbers (no public benchmarks). MEDIUM for the heuristic singing/spoken segmentation strategy (SING-03 — feasible from existing toolkit but unbenchmarked). --- ## Summary -Both spike URLs are **real, live, and the intended artifacts**. The "VoiceStudio" name is not overloaded in the wild — both `Serveurperso/OmniVoice-GGUF` and `ModelsLab/omnivoice-singing` are direct descendants of `k2-fsa/OmniVoice` (the same upstream model VoiceStudio already ships as its default cloning engine). License chain is clean: Qwen3-0.6B-Base → k2-fsa/OmniVoice (Apache-2.0) → both downstream variants (Apache-2.0). Both use the **same Higgs Audio v2 codec at 24 kHz mono**, the same Qwen3-0.6B language model backbone, and the same overall architecture — they differ only in (a) quantization + runtime (GGUF/`omnivoice.cpp`) and (b) finetune dataset + emotion/singing control tags (ModelsLab). +Both spike URLs are **real, live, and the intended artifacts**. The "VoiceStudio" name is not overloaded in the wild — both `Serveurperso/OmniVoice-GGUF` and `ModelsLab/omnivoice-singing` are direct descendants of `k2-fsa/OmniVoice` (the same upstream model VoiceStudio already ships as its default cloning engine). Their shared lineage does not establish commercial permission; see the licence correction above. Both use the **same Higgs Audio v2 codec at 24 kHz mono**, the same Qwen3-0.6B language model backbone, and the same overall architecture — they differ only in (a) quantization + runtime (GGUF/`omnivoice.cpp`) and (b) finetune dataset + emotion/singing control tags (ModelsLab). **The framing changes once that's confirmed.** SPIKE-01 is not "adopt a new engine" — it is "ship a quantized runtime variant of the engine already inside VoiceStudio, selectable by hardware probe." SPIKE-02 is not "adopt a new engine architecture" — it is "ship a domain-adapted finetune of the same VoiceStudio model with singing-mode tags, callable through the existing `VoiceStudioBackend` API surface with a different `from_pretrained` ID." **Primary recommendations:** -- **SPIKE-01 (OmniVoice-GGUF): GO**, conditional on a Phase 4-internal Apple Silicon `buildmetal.sh` smoke (no published Metal build script visible in `omnivoice.cpp/README.md`; only Vulkan and CPU are documented). The integration shape is `SubprocessBackend` wrapping the `omnivoice-tts` C++ CLI, with quant selected by a `detect_capabilities()` hardware probe. +- **SPIKE-01 (OmniVoice-GGUF): technical GO**, conditional on a Phase 4-internal Apple Silicon `buildmetal.sh` smoke (no published Metal build script visible in `omnivoice.cpp/README.md`; only Vulkan and CPU are documented). The integration shape is `SubprocessBackend` wrapping the `omnivoice-tts` C++ CLI, with quant selected by a `detect_capabilities()` hardware probe. - **SPIKE-02 (omnivoice-singing): GO with reduced scope**, treating it as a **second `from_pretrained` ID against the existing `VoiceStudioBackend`** rather than a new backend class — it is the same `omnivoice` PyPI library, same `transformers` pipeline, same model interface. The dubbing-pipeline "singing mode" toggle (SING-02) and Demucs vocal-stem routing (SING-03) remain real work, but they're pipeline integration, not engine integration. Net: 5 SING-* requirements stay in scope; SING-01's framing simplifies. **Phase 2 dependency:** Both engines build on the `SubprocessBackend` primitive from Phase 2. Phase 4 cannot finalize PLAN.md until Phase 2 RESEARCH.md exists and confirms the subprocess + venv + `mp.get_context("spawn")` + `HF_HOME` inheritance contract. **Capture this as a planner gate, not as research blocked-on-Phase-2** — research can proceed using `SubprocessBackend` as a stable contract (the ROADMAP and SUMMARY both define it). PLAN.md cannot reference its internals until Phase 2 RESEARCH lands. @@ -90,8 +110,8 @@ Both spike URLs are **real, live, and the intended artifacts**. The "VoiceStudio # Confirmed live on HuggingFace (verified 2026-05-18 via WebFetch): # Serveurperso/OmniVoice-GGUF — 10,603 downloads last month, 4 quant variants -# ModelsLab/omnivoice-singing — 1,053 downloads last month, Apache-2.0 -# k2-fsa/OmniVoice (upstream) — Apache-2.0, arXiv 2604.00688 +# ModelsLab/omnivoice-singing — historical download count; commercial rights unresolved +# k2-fsa/OmniVoice (upstream) — separate code/weight terms; arXiv 2604.00688 # Confirmed live on GitHub (verified 2026-05-18 via WebFetch): # ServeurpersoCom/omnivoice.cpp — MIT, 38 stars, 59 commits on master, 6 open issues @@ -324,7 +344,7 @@ class VoiceStudioGGUFBackend(TTSBackend): | Problem | Don't Build | Use Instead | Why | |---------|-------------|-------------|-----| | GGUF inference | A custom `llama.cpp` fork that adds the `omnivoice-lm` arch | `omnivoice.cpp` binary (upstream, MIT, by the quant author) | The author of the quants is the author of the runtime. Forking llama.cpp ourselves would add a permanent maintenance burden and zero capability over upstream. | -| Singing voice cloning model architecture | A trained-from-scratch singing model | `ModelsLab/omnivoice-singing` (same arch, finetuned) | The finetune already exists, license is clean, runtime is what we already ship. | +| Singing voice cloning model architecture | A trained-from-scratch singing model | `ModelsLab/omnivoice-singing` (same arch, finetuned) | The finetune exists and reuses the current runtime; commercial rights remain unresolved. | | Vocal stem isolation for singing-mode routing | A custom source-separation network | Demucs (already in dubbing pipeline) | Already a dep, already integrated, results are good enough for routing decisions. | | Speech-vs-singing classifier | A trained classification model for SING-03 v0.3 | Pitch-stability + energy heuristic on Demucs vocal stem; per-segment user override in UI | REQUIREMENTS.md already defers model-based classifier to v2. Heuristic is sufficient with override. | | Hardware probe / VRAM bucketing | A new GPU-detection library | Extend the existing `backend/services/gpu_sandbox.py` probe | Already detects CUDA/MPS/ROCm/CPU; just adds VRAM bucketing on top. | @@ -637,7 +657,7 @@ _REGISTRY.update({ | SING-01 | `VoiceStudioSingingBackend` loads, generates 1s with `[singing]` tag auto-injected | unit | `pytest tests/engines/test_omnivoice_singing.py::test_generate_with_auto_tag -x` | ❌ Wave 0 | | SING-02 | Dubbing pipeline routes vocal stem → singing engine when toggle is on; instrumental preserved | integration | `pytest tests/services/test_dub_pipeline_singing.py::test_singing_mode_preserves_instrumental -x` | ❌ Wave 0 | | SING-03 | Segment detector returns valid `Segment` list with kind + confidence on a known mixed clip | unit | `pytest tests/services/test_segment_detector.py::test_mixed_clip_routing -x` | ❌ Wave 0 | -| SING-04 | License surfacing endpoint returns Apache-2.0 + ModelsLab URL; first-use acceptance is gated | unit + UI | `pytest tests/engines/test_omnivoice_singing.py::test_license_gate -x` | ❌ Wave 0 | +| SING-04 | License surfacing must distinguish code, weights and upstream restrictions; acceptance/enforcement remains to be implemented | unit + UI | `pytest tests/engines/test_omnivoice_singing.py::test_license_gate -x` | ❌ Wave 0 | | SING-05 | 30s mixed speech+singing clip dubs end-to-end; both segments intelligible, instrumental preserved | smoke | `scripts/smoke-singing.sh tests/fixtures/mixed-30s.wav` | ❌ Wave 0 | ### Sampling Rate @@ -689,7 +709,7 @@ _REGISTRY.update({ ### How this research executed the SPIKE protocol -1. **Web-fetch model cards** [DONE 2026-05-18]: Confirmed both `Serveurperso/OmniVoice-GGUF` and `ModelsLab/omnivoice-singing` exist, are public, are descendants of `k2-fsa/OmniVoice`, and use compatible licenses. Verified the upstream chain via `huggingface.co/k2-fsa/OmniVoice`. +1. **Web-fetch model cards** [DONE 2026-05-18]: Confirmed both `Serveurperso/OmniVoice-GGUF` and `ModelsLab/omnivoice-singing` exist, are public, are descendants of `k2-fsa/OmniVoice`, with shared upstream lineage. The original licence-compatibility conclusion is withdrawn; see the correction above and `huggingface.co/k2-fsa/OmniVoice`. 2. **Web-fetch runtime repo** [DONE 2026-05-18]: Confirmed `github.com/ServeurpersoCom/omnivoice.cpp` is the only runtime for the quants; documented MIT license, build script set, CLI invocation pattern, and maintenance state (38 stars, 59 commits, 6 open issues). 3. **PyPI package verification** [DONE 2026-05-18]: Confirmed `omnivoice` 0.1.5 on PyPI (2026-04-28, Apache-2.0) — already a project dep, no new dependency for SPIKE-02. 4. **Architecture compatibility check** [DONE]: Both downstream models share the same architecture as the project's existing default `VoiceStudioBackend` (Qwen3-0.6B + Higgs Audio v2 codec, 24 kHz mono). This is the load-bearing finding that re-frames both spikes from "new engines" to "variants of the engine already shipping." @@ -707,12 +727,14 @@ _REGISTRY.update({ ## GO / NO-GO Recommendations -### SPIKE-01 (Serveurperso/OmniVoice-GGUF): **GO** ✓ +### SPIKE-01 (Serveurperso/OmniVoice-GGUF): **TECHNICAL GO** + +This establishes technical feasibility only. Free-app acceptance also requires the rights gate in [engine acceptance](../engine-acceptance.md), including owner approval and first-use disclosure for restricted engines. **Rationale:** -- ✓ Model card, runtime repo, license, lineage all verified live (2026-05-18). +- Historical model/runtime/lineage review dated 2026-05-18; commercial clearance superseded above. - ✓ Same underlying model as VoiceStudio's existing default — this is *quantization of what we already ship*, not a new engine architecture. Worst case it just doesn't beat the in-process Python path on a given hardware class, and we keep that path as the fallback. -- ✓ License chain clean: Apache-2.0 (model) + MIT (runtime). +- Commercial model clearance withdrawn; runtime and weight terms require separate records. - ✓ Cross-platform via CUDA / Vulkan / Metal / CPU per `omnivoice.cpp` build matrix [with **Assumption A1** caveat — Metal build script not in published README, must validate Wave 1]. - ✓ Hardware adaptation is a natural fit for the user-stated value of "first-run that actually works" on a wide range of hardware — Q4_K_M ~659 MB VRAM gets the engine running on 4 GB GPUs that today fall back to CPU on the in-process Python path. @@ -725,9 +747,9 @@ _REGISTRY.update({ ### SPIKE-02 (ModelsLab/omnivoice-singing): **GO** with reduced scope ✓ **Rationale:** -- ✓ Model card, license, runtime path verified live (2026-05-18). +- Historical model/runtime review dated 2026-05-18; commercial clearance superseded above. - ✓ Same `omnivoice` PyPI library, same architecture, same codec — load-bearing finding: this is *not* a new engine, it's a new model ID consumed by the engine class we already have. `VoiceStudioSingingBackend` is a ≤30-line subclass. -- ✓ License clean: Apache-2.0 with documented training-data downstream compliance (training datasets carry CC BY-NC-SA / ODbL constraints, which propagate to commercial *training* but not to *use* of the model under Apache-2.0). +- The former training-versus-use licence conclusion is withdrawn. The singing model requires upstream clarification before commercial approval. - ✓ Hardware: same footprint as existing VoiceStudio; runs on existing-engine-compatible hardware. **Scope reductions captured by this research:** @@ -763,20 +785,20 @@ Two files, written by the planner after Phase 4 PLAN.md is locked, using this re ## Context -VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (Apache-2.0, 0.6B Qwen3 backbone, Higgs Audio v2 codec) as its default voice-cloning engine via `backend/services/tts_backend.py:VoiceStudioBackend`. The Python in-process path requires PyTorch + CUDA / MPS / CPU. +VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (separate code/weight terms; 0.6B Qwen3 backbone, Higgs Audio v2 codec) as its default voice-cloning engine via `backend/services/tts_backend.py:VoiceStudioBackend`. The Python in-process path requires PyTorch + CUDA / MPS / CPU. `Serveurperso/OmniVoice-GGUF` publishes 4 quantizations of the same model (Q4_K_M / Q8_0 / BF16 / F32) consumable through the MIT-licensed `omnivoice.cpp` runtime (a custom GGML-based C++ inference binary). This decision is whether to integrate the GGUF engine as a hardware-adaptive default with overridable fallback. ## Decision -**GO** — integrate per GGUF-01..06. +**Technical GO only** — integrate per GGUF-01..06; engine acceptance remains conditional on the rights gate in `docs/engine-acceptance.md`. ## Consequences **Positive:** - 4 GB-VRAM GPUs (currently falling back to CPU on the in-process path) get GPU-backed cloning via Q4_K_M. - Smaller VRAM footprint = stays out of the way of other engines when users run multiple in one session. -- License chain unchanged (Apache-2.0 model + MIT runtime). +- Model licensing needs separate review from the runtime licence; see the correction above. **Negative / risk:** - Adds a maintained-by-others C++ runtime to the dependency graph (`omnivoice.cpp`, 38 stars at decision time). @@ -793,7 +815,7 @@ VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (Apache-2.0, 0.6B Qwen3 backbone, Hi - `.planning/phases/04-adaptive-specialty-engines-spike-first/04-RESEARCH.md` (this research) - https://huggingface.co/Serveurperso/OmniVoice-GGUF (verified 2026-05-18) - https://github.com/ServeurpersoCom/omnivoice.cpp (verified 2026-05-18) -- https://huggingface.co/k2-fsa/OmniVoice (upstream, Apache-2.0) +- https://huggingface.co/k2-fsa/OmniVoice (upstream; separate code/weight terms) ``` ### `.planning/decisions/SPIKE-02-singing.md` @@ -808,7 +830,7 @@ VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (Apache-2.0, 0.6B Qwen3 backbone, Hi ## Context -`ModelsLab/omnivoice-singing` is a finetune of `k2-fsa/OmniVoice` (same Apache-2.0, same Qwen3-0.6B backbone, same Higgs Audio v2 codec, same `omnivoice` PyPI library) trained on additional singing + emotion-tagged data. Activated by a `[singing]` text control tag at generation time. +`ModelsLab/omnivoice-singing` is a finetune of `k2-fsa/OmniVoice` (commercial rights unresolved; same Qwen3-0.6B backbone, same Higgs Audio v2 codec, same `omnivoice` PyPI library) trained on additional singing + emotion-tagged data. Activated by a `[singing]` text control tag at generation time. VoiceStudio's dubbing pipeline currently routes vocal stems (via Demucs) through the default TTS engine, which produces speech output even on sung source material. This decision is whether to integrate the singing finetune as a routed alternative for sung segments. diff --git a/docs/adr/SPIKE-01-gguf.md b/docs/adr/SPIKE-01-gguf.md index a20bf8f30..07604b7a4 100644 --- a/docs/adr/SPIKE-01-gguf.md +++ b/docs/adr/SPIKE-01-gguf.md @@ -3,14 +3,19 @@ # SPIKE-01: Adopt `Serveurperso/OmniVoice-GGUF` as hardware-adaptive default cloning engine -**Status:** Proposed (research-supported) — Wave 1 build/smoke flips to Accepted in Task 3 +> **Licence correction (2026-10-03):** the original commercial-compatibility +> conclusion is withdrawn. See the [source review](SPIKE-01-gguf-research.md#licence-correction-2026-10-03) +> for separate code, weight, and downstream terms. This engineering decision +> does not approve the model for Pro. + +**Status:** Proposed — Wave 1 build/smoke establishes technical feasibility only. Acceptance also requires the rights gate in [engine acceptance](../engine-acceptance.md), including owner approval and first-use disclosure for restricted free-app engines. **Date:** 2026-05-18 (updated 2026-05-20 with pinned SHAs) **Decision-makers:** [maintainer] **Related:** ROADMAP Phase 4; REQUIREMENTS GGUF-01..06; `.planning/phases/04-adaptive-specialty-engines-spike-first/04-RESEARCH.md` ## Context -VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (Apache-2.0, 0.6B Qwen3 backbone, Higgs Audio v2 codec at 24 kHz mono) as its default voice-cloning engine via `backend/services/tts_backend.py:VoiceStudioBackend`. The Python in-process path requires PyTorch + CUDA / MPS / CPU and on 4 GB-VRAM GPUs falls back to CPU inference. +VoiceStudio v0.2.7 ships `k2-fsa/OmniVoice` (separate code/weight terms; 0.6B Qwen3 backbone, Higgs Audio v2 codec at 24 kHz mono) as its default voice-cloning engine via `backend/services/tts_backend.py:VoiceStudioBackend`. The Python in-process path requires PyTorch + CUDA / MPS / CPU and on 4 GB-VRAM GPUs falls back to CPU inference. `Serveurperso/OmniVoice-GGUF` (HuggingFace, 10,960 downloads/month, verified 2026-05-20) publishes 4 quantizations of the same upstream model — Q4_K_M (~659 MB VRAM), Q8_0 (~945 MB, recommended balance), BF16 (~1.6 GB), F32 (~3.2 GB) — consumable through the MIT-licensed `omnivoice.cpp` runtime (`github.com/ServeurpersoCom/omnivoice.cpp`, 42 stars, 59 commits, 6 open issues at the pinned SHA). The quants use a custom `omnivoice-lm` architecture (confirmed via HF API `gguf.architecture`) and do **not** load in vanilla llama.cpp. @@ -18,7 +23,7 @@ This decision is whether to integrate the GGUF engine as a hardware-adaptive def ## Decision -**GO** — integrate per GGUF-01..06. +**Technical GO only** — integrate per GGUF-01..06. Do not mark the engine Accepted until the rights gate above passes. The integration shape is `VoiceStudioGGUFBackend(TTSBackend)` wrapping Phase 2's `SubprocessBackend`, which spawns a bundled per-platform `omnivoice-tts` binary built from a pinned `omnivoice.cpp` commit SHA. Quant selection is driven by a `detect_capabilities()` extension of `backend/services/gpu_sandbox.py` mapping `(compute_class) → quant filename` via shippable `quant_map.json`. On hardware where probe + load succeed, GGUF becomes the default cloning engine; on any failure the existing in-process `VoiceStudioBackend` is the fallback. @@ -27,7 +32,7 @@ The integration shape is `VoiceStudioGGUFBackend(TTSBackend)` wrapping Phase 2's | Question | Verdict | Evidence | |----------|---------|----------| | Is the model the intended artifact (a quantization of `k2-fsa/OmniVoice`, not an overloaded "VoiceStudio" name)? | YES | HF model card explicitly chains: `Qwen/Qwen3-0.6B-Base → Qwen/Qwen3-0.6B → k2-fsa/OmniVoice → Serveurperso/OmniVoice-GGUF`. `base_model:k2-fsa/OmniVoice` and `base_model:quantized:k2-fsa/OmniVoice` tags present in the API response. | -| License compatible with v0.3.x ship? | YES — Apache-2.0 (model) + MIT (runtime) | Both verified via HF model card + GitHub README. Same Apache-2.0 chain as the upstream model already shipping in v0.2.7. | +| License compatible with commercial use? | Original approval withdrawn | See the licence correction above; runtime and model terms differ. | | Runtime: llama.cpp / candle / custom? | CUSTOM (`omnivoice.cpp`, MIT) — does NOT load in vanilla llama.cpp | `gguf.architecture = "omnivoice-lm"` from HF API; README states "GGUF weights for omnivoice.cpp, a C++17/GGML port of VoiceStudio". | | Quant variants and footprints? | 4 quants × 2 files each (base + tokenizer): Q4_K_M (659 MB), Q8_0 (945 MB), BF16 (1.60 GB), F32 (3.19 GB) | HF `siblings` list confirms all 8 files; sizes from model card table. | | Cross-platform runtime fit? | Linux + Windows + macOS Intel YES via documented build scripts; macOS Apple Silicon Metal YES (builds clean via `cmake -DGGML_METAL=ON` at pinned SHA, #2105) | `buildcpu.sh`, `buildcuda.sh`, `buildvulkan.sh`, `buildall.sh` listed; Metal builds cleanly via `cmake -DGGML_METAL=ON` in CI and locally (#2105). | @@ -45,7 +50,7 @@ Both SHAs are mirrored in `backend/engines/omnivoice_gguf/quant_map.json` `_meta **Positive:** - 4 GB-VRAM GPUs (currently falling back to CPU on the in-process path) get GPU-backed cloning via Q4_K_M. - Smaller VRAM footprint = stays out of the way of other engines when users run multiple in one session. -- License chain unchanged (Apache-2.0 model + MIT runtime). +- Runtime integration does not establish commercial model rights; see the correction above. - Same underlying model as what already ships — worst case it ties the in-process path on a given hardware class and we keep that path as the fallback. **Negative / risk:** @@ -69,5 +74,5 @@ Both SHAs are mirrored in `backend/engines/omnivoice_gguf/quant_map.json` `_meta - https://huggingface.co/api/models/Serveurperso/OmniVoice-GGUF (SHA `361609388ae572a820d085185bbbe2a2aac4b30e`, 2026-05-20) - https://github.com/ServeurpersoCom/omnivoice.cpp (verified 2026-05-18 / re-verified 2026-05-20) - https://api.github.com/repos/ServeurpersoCom/omnivoice.cpp/commits/master (SHA `886fc079838ca7400cb2b42b36e2a65aa1daabe8`, 2026-05-20) -- https://huggingface.co/k2-fsa/OmniVoice (upstream, Apache-2.0) +- https://huggingface.co/k2-fsa/OmniVoice (upstream; separate code/weight terms) - `backend/services/tts_backend.py` (existing `VoiceStudioBackend` reference) diff --git a/docs/adr/SPIKE-02-singing.md b/docs/adr/SPIKE-02-singing.md index 37e6b246c..bece24d86 100644 --- a/docs/adr/SPIKE-02-singing.md +++ b/docs/adr/SPIKE-02-singing.md @@ -3,6 +3,11 @@ # SPIKE-02: Adopt `ModelsLab/omnivoice-singing` as singing variant of the existing engine +> **Licence correction (2026-10-03):** the original commercial-compatibility +> conclusion is withdrawn. See the [source review](SPIKE-01-gguf-research.md#licence-correction-2026-10-03) +> for separate code, weight, and downstream terms. This engineering decision +> does not approve the model for Pro. + **Status:** ⚠️ **SUPERSEDED (2026-06-14)** by `specs/006-dubbing-singing-mode/` (spec tree removed 2026-07-12 — feature shipped; see git history) **Date:** 2026-05-18 (superseded 2026-06-14) **Decision-makers:** [maintainer] @@ -17,7 +22,7 @@ ## Context -`ModelsLab/omnivoice-singing` (HuggingFace, 1,053 downloads/month, verified 2026-05-18) is a finetune of `k2-fsa/OmniVoice` — same Apache-2.0 license, same Qwen3-0.6B backbone, same Higgs Audio v2 codec at 24 kHz mono, same `omnivoice` PyPI library (0.1.5, 2026-04-28) already shipping in VoiceStudio v0.2.7. Trained on additional singing + emotion-tagged data and activated by a `[singing]` text control tag at generation time. +`ModelsLab/omnivoice-singing` (HuggingFace, 1,053 downloads/month, verified 2026-05-18) is a finetune of `k2-fsa/OmniVoice` — unresolved commercial rights, same Qwen3-0.6B backbone, same Higgs Audio v2 codec at 24 kHz mono, same `omnivoice` PyPI library (0.1.5, 2026-04-28) already shipping in VoiceStudio v0.2.7. Trained on additional singing + emotion-tagged data and activated by a `[singing]` text control tag at generation time. VoiceStudio's existing `dub_pipeline.py` runs Demucs to split source audio into vocal and instrumental stems and routes the vocal stem through the default TTS engine. Today this produces speech-like output even on sung source material, which is one of the loudest user complaints when dubbing music-adjacent content. diff --git a/docs/electron-playback.md b/docs/electron-playback.md index 44e6c9866..18cab7aca 100644 --- a/docs/electron-playback.md +++ b/docs/electron-playback.md @@ -16,8 +16,7 @@ The test uses synthetic WAV/profile/generation responses and never saves user da Set `PLAYWRIGHT_CHANNEL` or `VOICESTUDIO_UI_URL` for another installed browser/server. HLS and DASH libraries are bundled and lazy-loaded. The shared provider includes -Vidstack's native audio/video, HLS, DASH, YouTube and Vimeo selection plus the -Remotion loader. Gallery search previews exercise the embedded YouTube provider; +Vidstack's native audio/video, HLS, DASH, YouTube and Vimeo selection. Gallery search previews exercise the embedded YouTube provider; Dubbing uses the custom Vidstack video controls for local and normalized URL imports. Native video MIME hints select a provider before its `