Skip to content

feat(studio): consolidate community fixes and composer UI - #2419

Merged
debpalash merged 96 commits into
mainfrom
consolidate/engine-matrix-ui-20260929
Sep 29, 2026
Merged

debpalash merged 96 commits into
mainfrom
consolidate/engine-matrix-ui-20260929

Conversation

@debpalash

@debpalash debpalash commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Summary

Consolidates 20 community PRs and completes the composer, sidebar and voice-card UI. Engine switches preserve active renders, and compressed audio remains usable without TorchCodec.

Changes

Type

  • New feature
  • Bug fix
  • Refactor
  • Documentation
  • Tests
  • CI / Build
  • Release prep

Testing

  • Local bun run check:electron: 1,143 Electron tests and 2,922 shared tests passed, one skipped; typecheck, desktop/web builds, packaging contract and all 21 locales passed.
  • Offline backend validation uses HF_HUB_OFFLINE=1 with a fresh empty HF_HUB_CACHE: targeted suites 735 passed/3 skipped; isolated backend suite 464 passed/8 skipped; later review regression suites 183 passed/15 skipped.
  • All 19 audio-loader tests pass, including real MP3/M4A/AAC/Opus decoding, bounded memory transport, compressed-input and decoded-output size limits, temporary-file cleanup, and disk headroom. Oversized output is rejected rather than returned truncated.
  • OuteTTS downmix/resampling regressions failed before the fix and pass afterward; affected engine/audio suites: 83 passed, 8 skipped.
  • All four browser smokes pass: playback, dubbing, longform layout, and language selection at desktop/compact/mobile widths. Frozen Bun installation and uv lock --check pass.
  • Four-platform packaging and fresh runtime installation passed on Windows, Linux, Apple Silicon and Intel Mac. The Intel smoke verifies remote setup on its unsupported local-runtime platform. Subsequent backend-only fixes bound compressed-audio staging and route OuteTTS reference preprocessing through the same decoder.
  • Final-head CI passed: 9,039 backend tests, 472 isolated backend tests, 1,144 Electron tests, 2,922 shared tests, browser workflows, and all four OS smokes. Backend results also include 32 skips, 8 expected failures and one non-strict unexpected pass.
  • CUDA and ROCm Docker builds and security checks passed on final head d223241a.

Earlier live MPS engine measurements remain in the handoff/history. These Windows checks do not claim new live synthesis for every optional model or a completed MPS ASR matrix. MeloTTS validates its documented manual g2p_en/NLTK prerequisites without downloading during generation.

Checklist

  • Tested locally, including fail-before/pass-after regressions.
  • Relevant documentation and contributor credits updated.
  • No local machine paths, logs or personal environment details committed.
  • Version files remain in sync; version 0.5.6 is unchanged.
  • Final smoke-matrix fixture checks green on macOS, Windows and Linux.

Review disposition: all concrete CodeRabbit findings and engine/export/backup Greptile findings are fixed. Credential-free server mode deliberately permits consumption and export recording; its history redacts host paths. Configuring a key or PIN rejects anonymous access. Native file operations retain separate authorization. The lazy import-cycle note has no import-time cycle. The sidebar remains collapsible on every platform; macOS uses the existing sidebar controls. The automatic public star refresh follows the explicit owner exception in CLAUDE.md. MCP voice creation uses the existing admitted-caller write boundary, matching clone_voice. Greptile withdrew its final truncation warning after the FFmpeg documentation, muxer code, and nine real size-boundary checks confirmed the guard.

Release cadence

Continuous-to-main: this merge joins rolling source and Docker :latest. Electron rehearsals publish artifacts only. No version bump, release tag, or stable release is requested; those remain owner-controlled under the release checklist.

aiapienthusiast and others added 30 commits September 24, 2026 20:14
Updated supported LLM providers list and added details for Cheaper Inference configuration.
Updated expected provider list to include 'cheaperinference' and added tests for its contract.
With Korean, Japanese or Chinese input methods, the Enter that confirms
a composition is delivered as a keydown with isComposing (or keyCode 229
on macOS). Several text-field handlers treated it as a submit, so
confirming a conversion also renamed a project, picked a language, added
a pronunciation rule, issued an inbound key, renamed a worker or saved an
MCP binding.

Add isImeComposing() in lib/ime.ts, guard the unguarded handlers with it,
and route the existing command palette and dub shortcut guards through
the same helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Agents connected over MCP could clone voices but not design them: the
design path (POST /design/describe + POST /profiles kind=design) had no
MCP tool, and generate_speech's instruct only restyles a resolved profile.

describe_voice previews the attribute mapping without saving; design_voice
saves a design profile and returns its profile_id, refusing descriptions
that map to no attribute.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- generate_speech omits language unless given, so a profile's saved
  language applies (#533); an explicit Auto still overrides it.
- design_voice waits a generation-sized deadline: the save renders the
  identity sample through the GPU queue.
- describe requests honor the configured MCP POST timeout.
- Drop design_voice's personality argument: generation never applies a
  stored personality, so it had no effect.
- Document Chinese dialects and the deferred sample render.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit updates README_CN.md to match the current Electron app positioning and install flow. It adds product positioning, badges, docs links, one-click install commands, source-run instructions, migration guidance for Tauri users, workspace and feature highlights, and refreshed sponsor/support sections.
torchaudio 2.9 routes torchaudio.load() through TorchCodec, which is not
in uv.lock — so on every uv sync source install, dub assembly raises
ImportError and every generate fails with dub_speech_missing even though
the TTS segments rendered fine. The save side of this backend shift is
already guarded in services/audio_io.py; this adds the matching fallback
for the load side, reading the rendered PCM WAV segments via soundfile.

Fixes #2378

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
The canvas has always drawn a condition with two source handles, the
palette offered it, and `onConnect` already carried the handle into the
document — but `compileWorkflow` rejected anything that forked, so the
step could only ever be placed, never run. It shipped labelled "local
draft". This makes it execute.

A condition compares the text reaching it against its phrase — contains,
is exactly, starts with, ends with — and sends the item down the `yes` or
`no` branch. Matching ignores capitals and surrounding blanks, and uses
`toLowerCase` rather than the locale-aware form deliberately: the same
draft has to take the same branch on every machine, and Turkish dotless i
would otherwise make a Turkish host miss a phrase a US host matches.

Compilation becomes a graph walk instead of a straight line. It records
the value kind each node is ENTERED with, so paths that re-merge must
agree on it — one branch speaking while the other passes text straight
through is refused up front rather than handing the shared tail a
different kind of value per item. That record also stops a shared tail
being re-validated once per branch, and the walk keeps rejecting cycles,
forks from anything but a condition, unlabelled or duplicated branch
edges, and stranded steps. Execution then walks the graph per item, so
one run legitimately sends one clip down `yes` and the next down `no`.

Conditions are re-evaluated on resume rather than recorded: the text they
test is itself checkpointed, so a resumed run re-reads the same value and
takes the same branch. Wiring joins the run signature, because with a
fork in the graph, moving an edge changes which steps an item passes
through without changing any of them.

`workflowRun.invalid_graph` told users branches were not allowed; it and
the merge rule are retranslated across all 21 locales.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Unauthenticated callers on a LAN could inject export history records
or read local filesystem paths exposed in export history. Added
`dependencies=[Depends(require_loopback)]` to both endpoints, matching
the pattern already used by DELETE /export/history/{id}.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CV9A88tyfnofgQu4Dyj16
The previous SELECT MAX(seq) + INSERT pattern had a read-then-write
gap: two concurrent callers could read the same MAX and produce
duplicate seq values. job_events has no UNIQUE(job_id, seq) constraint,
so the duplicates silently landed in the DB and confused SSE replay.

Replace with a single atomic INSERT … SELECT … RETURNING that computes
and writes the next seq in one SQLite statement. No gap, no duplicate.
Requires SQLite ≥ 3.35 (RETURNING support, released 2021-03-12);
the project already requires Python ≥ 3.11 whose bundled SQLite is
≥ 3.39 on every supported platform.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016CV9A88tyfnofgQu4Dyj16
Hotkey dictation could not pass Whisper an initial_prompt, so names,
jargon and code-switched terms were mis-heard. Add an optional
`dictation.prompt` pref (GET/POST /dictation/prefs, Settings -> Dictation
shortcut) and hand it to the capture engine as initial_prompt on the live
socket and REST /transcribe (not `reference` mode). Engines whose
transcribe() does not declare it never receive it, via the signature
filter now shared with the OpenAI-compatible route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2FE3m1Fw5KJ7a9Uzq9DEN
…ngines

Local review follow-ups:
- REST /transcribe applies the hint only with dictation=true (the capture
  widget's fallback sends it); file transcription and MCP/CLI stay unbiased.
- The silent-sherpa-model rescue decodes without the hint on both the
  socket and REST: its text is the demotion evidence and a prompted
  Whisper can echo the prompt on noise.
- Bound the stored prompt on read as well as write.
- Settings copy names the engines that take it (Faster Whisper, MLX
  Whisper, OpenAI-compatible) and says it is sent to an OpenAI-compatible
  server; 21 locales.
- The field is locked while saving and follows the saved value again once
  an edit is reverted.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2FE3m1Fw5KJ7a9Uzq9DEN
PyTorch ships no macOS x86_64 wheels so the dependency set can never resolve there (#889). Preflight had no platform check, so setup burned a multi-GB uv sync before dying on the resolver error (#2365). A platform fail check now blocks Continue on the wizard's first step with remote-backend guidance.
_CURRENCY_RE's lookahead (?![\d.,]) kept $1,000 and $5.5 untouched but also rejected a plain period or comma, so "It costs $5." and "Pay $5, please" reached the engine as digits while "It costs $5" was spoken. It now rejects only a digit or a separator followed by a digit, the guard the integer and decimal rules use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…2390)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…spaces

apply_lexicon wraps a key in \b on each word-character edge. Japanese, Chinese and Thai write words without spaces, so a key there sits between more word characters and \b never matches: an entry applied only to a line that was the key alone. An edge in one of those scripts (dub_qc's _NO_SPACE_SCRIPT set) now gets no boundary; spaced scripts keep theirs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…spaces (#2392)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…names

The prompt filter matches on transcribe() signatures, so a real engine
collapsing to **kwargs would drop the hint silently while every stub-based
test stays green.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2FE3m1Fw5KJ7a9Uzq9DEN
_TextExtractor kept only the first text node of the first heading, so "<h1>Chapter <em>One</em></h1>" named the chapter "Chapter", "<span>3</span> The Calm" named it "3", and the rest of the heading was dropped from both title and body. The heading's text is now collected until it closes and whitespace-collapsed, with a <br> inside it separating words.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e letter

U+3005 (the iteration mark in words such as the Japanese for "various"), U+3006, U+3007 and U+303B are letters of the no-space scripts but were missing from dub_qc's set, so a lexicon key ending in one still got a trailing word boundary and missed mid-sentence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.