Streaming and high-throughput DiffSinger synthesis on a laptop CPU. No GPU required.
DiffSinger-ASM runs the complete acoustic-to-waveform pipeline in C and handwritten x86-64 assembly. It keeps separate execution and performance contracts for low-latency streaming and large-block batch rendering because a single long-audio RTF cannot describe both workloads honestly.
On an Intel Core i5-13420H, the current 32-frame streaming profile keeps every measured region below RTF 1 with zero deadline misses. The independent 384-frame batch profile renders 31.208 seconds of audio in a median 22.031 seconds, an aggregate RTF of 0.706, entirely on the CPU.
Exported DiffSinger acoustic and NSF-HiFiGAN ONNX models are compiled offline into memory-mapped native bundles. ONNX Runtime helps pack and validate a voicebank, but the production inference path needs no ONNX Runtime, PyTorch, oneDNN, OpenVINO, MKL, BLAS, or GPU.
The product-facing libdsasm.so keeps one native engine per singer in process,
reuses its model mappings, worker pool, and inference buffers, and publishes
completed PCM blocks directly to the playback mixer. OpenUtau can therefore
start playback from the first vocoder bucket instead of waiting for a complete
phrase WAV.
Important
This remains an experimental, CPU-specific runtime, not a standalone singing editor. The v0.2 provider package contains the native ABI and offline model tool required by the matching OpenUtau integration, but it is not part of an upstream OpenUtau release. Bring your own compatible exported voicebank.
Streaming and batch results are separate champions. For streaming, worst-region RTF below 1 is a real-time feasibility boundary: every block must finish before its playable audio is exhausted. It is not the score to minimize after the service gate passes. The optimization objective then moves to CPU-RTF and concurrent-track capacity while preserving zero deadline misses. Batch rendering continues to optimize complete E2E wall-time RTF. Results are never promoted by comparing one architecture with the other.
| Contract | Current champion |
|---|---|
| Streaming workload | 32 frames, 25 measured regions after 10 warm-ups |
| Streaming service gate | 0.996841 largest worst-region RTF across three runs; required < 1 |
| Streaming CPU-RTF | 2.391249 maximum across three runs |
| Streaming deadlines | 0 misses in every run |
| Streaming waveform quality | cosine 0.999292, SNR 28.49 dB |
| Batch workload | 7 x 384-frame regions after 2 warm-ups |
| Batch generated audio | 31207.619 ms |
| Batch median E2E latency | 22031.032 ms |
| Batch median aggregate RTF | 0.705950 |
| Batch p90 / worst latency | 24730.666 / 25563.829 ms |
| Batch waveform quality | cosine 0.999033, SNR 27.13 dB |
Both profiles were measured on an Intel Core i5-13420H at PL1 45 W with the performance platform and EPP policies. Streaming uses four workers to reduce per-track CPU load; batch uses eight P-core/SMT workers for throughput. The batch champion improved its paired control by 8.02% while improving p90 and worst latency.
See the streaming policy and evidence, the batch policy and evidence, and the tracked batch raw samples. Historical M55/M58 long-audio results remain in the performance record, but are not current streaming or fixed-block champions. Results on other CPUs and voicebanks will vary.
Interactive playback and offline rendering put pressure on different parts of the system. One block size cannot minimize response time and maximize sustained throughput at the same time, so DiffSinger-ASM supplies two independently tuned engines:
| Mode | Used for | Execution shape | Performance focus |
|---|---|---|---|
| Real-time streaming | Playback while editing | 32-frame vocoder buckets, 8-frame overlap, 4 workers, progressive PCM callbacks | Service constraint: worst-region RTF below 1 with zero deadline misses. Optimization after that: CPU-RTF and concurrent-track capacity |
| Block batch | Pre-render, mixdown, and export | 384-frame buckets, no overlap, 8 workers, complete blocks | End-to-end throughput, aggregate RTF, p90 and worst latency, and run-to-run stability |
Small streaming blocks bound the time before the mixer receives audio, but they repeat scheduling, synchronization, and overlap work more often. Large batch blocks amortize that overhead and keep more arithmetic in flight, but waiting for a large block would make interactive playback feel unresponsive. Separate worker counts and kernels let each workload optimize the cost that its user actually notices.
The streaming metrics form a staged optimization gradient. Before the service gate is reached, reducing the largest worst-region RTF is necessary. After it is reached, a lower RTF does not by itself make a better streaming engine: a candidate advances by reducing CPU-RTF and increasing usable track capacity while every region remains below 1 and every deadline still passes. Batch RTF has a different meaning because batch users are waiting for the complete job.
The modes are execution choices, not quality levels. Both pass the same model correctness and waveform quality gates, and both declare output compatibility revision 1. Once either mode completes a phrase, OpenUtau may reuse that canonical PCM for later playback, pre-rendering, mixdown, or export.
- Each product mode has its own target. Streaming protects deadlines and CPU capacity; batch rendering maximizes complete E2E throughput.
- The hot path stays small. Runtime dependencies are only libc, libm, and pthread, with no framework startup, graph planner, or provider dispatch.
- Kernels match deployed shapes. Handwritten AVX2/FMA code handles FP32 operators; asymmetric AVX-VNNI accelerates quality-gated vocoder stages.
- Models are prepared ahead of time. Fixed shapes, packed weights, fused residual operations, and memory-mapped bundles move work out of inference.
- Persistent workers keep the CPU busy. Scheduling is tuned for the performance cores and avoids rebuilding execution state for every request.
- Speed never bypasses quality. Releases must pass waveform parity, stability, p90, and worst-case gates before they are promoted.
This project is an independent native inference runtime for deployment models exported by the OpenVPI-maintained DiffSinger project. It does not train voicebanks and does not replace the upstream variance model or score frontend. The integrated OpenUtau renderer constructs those inputs and selects this runtime as an alternative to ONNX Runtime.
DiffSinger architecture overview from OpenVPI/DiffSinger, licensed under Apache-2.0.
Use the upstream ecosystem to prepare a model and musical inputs, then use DiffSinger-ASM for the acoustic and vocoder execution stages:
- DiffSinger user guidance and Getting Started
- MakeDiffSinger for dataset and voicebank preparation
- OpenUtau for production-oriented score editing and synthesis workflows
- DiffSinger paper and OpenVPI implementation
DiffSinger-ASM is used through the matching OpenUtau build. Normal users do not need a compiler, Python, ONNX Runtime, or command-line setup.
- Download and extract the Linux x64 build from AntheaLaffy's OpenUtau releases.
- Download and extract the latest
DiffSinger-ASM provider.
In OpenUtau, open Preferences > Rendering, select the provider's
lib/libdsasm.sounder DiffSinger ASM native library, and confirm that its status is ready. - On a DiffSinger track, select DIFFSINGER-ASM as the renderer.
- Open Tools > Singers, select the voicebank, and choose Convert current model. OpenUtau retains the ONNX source and publishes the converted model only after validation succeeds.
The singer is ready when its status changes to ASM model installed. Complete pre-renders and completed real-time renders share the canonical PCM cache, so playback reuses audio OpenUtau has already rendered.
The provider currently requires Linux x86-64 with AVX2 and FMA. Conversion adds about 550 MB per singer and can take several minutes. Unsupported model graphs are reported without modifying the source voicebank.
The release package already includes its offline dependencies. A source
checkout needs a C toolchain, GNU Make, binutils, CPython 3.12, and uv:
git clone https://github.com/Asmory/DiffSinger-ASM.git
cd DiffSinger-ASM
python3.12 -m venv .venv
. .venv/bin/activate
python -m pip install -r tools/model-tool-requirements.inInstall PyTorch in this environment when working on checkpoint importers or
golden/parity validation. Install Linux perf when profiling. See the
developer guide for dependency boundaries, repository
layout, agent and tool routing, experiment scheduling, validation gates,
benchmark rules, packaging, and OpenUtau integration. It translates the core
rules in AGENTS.md into a workflow intended for both human contributors and
coding agents.
make -j"$(nproc)" engine-check
make -j"$(nproc)" package VERSION=v0.2.2This produces build/libdsasm.so and a self-contained archive under
release/. The package embeds Python, NumPy, ONNX, and ONNX Runtime for offline
conversion; native inference does not load them.
Check CPU features when testing a local build:
grep -m1 -oE 'avx2|fma|avx_vnni' /proc/cpuinfo | sort -uInspect, plan, convert into caller-owned staging, and validate without loading the native runtime:
package=/absolute/path/to/diffsinger-asm-v0.2.2-linux-x86_64
singer=/absolute/path/to/singer
"$package/bin/dsasm-model-tool" inspect \
--protocol 1 --singer-root "$singer" --json
"$package/bin/dsasm-model-tool" plan \
--protocol 1 --singer-root "$singer" --json
"$package/bin/dsasm-model-tool" convert \
--protocol 1 --singer-root "$singer" \
--expected-source-fingerprint SHA256_FROM_PLAN \
--staging /path/on/singer/filesystem/staging \
--work /path/to/reusable/work --jsonl
"$package/bin/dsasm-model-tool" validate \
--protocol 1 --bundle /path/on/singer/filesystem/staging --jsonConversion always produces the 32- and 384-frame product buckets and writes
bundle.json only after offline validation succeeds. Staging must be empty and
must not overlap the reusable work directory. SIGINT cancels and reaps the
active packer process group before the tool returns. The tool never publishes
current.json; OpenUtau owns the atomic generation commit.
The direct packer commands below are retained for development and standalone
debugging. Product integrations should use dsasm-model-tool so compatibility,
fingerprints, estimates, reason codes, and validation stay on one versioned
boundary.
The standalone acoustic and vocoder executables remain available for model inspection, parity checks, and debugging:
make -j"$(nproc)" build/dsasm-acoustic build/dsasm-vocoder-m40Pack the acoustic model into the singer's default ASM directory:
python tools/pack_acoustic_onnx_m25.py /path/to/voicebank/acoustic.onnx \
--model-dir /path/to/voicebank \
--out /path/to/singer/dsasm/acoustic
build/dsasm-acoustic inspect /path/to/singer/dsasm/acousticCompile the two product vocoder buckets. The filenames are part of the engine contract; each graph remains fixed-shape internally:
mkdir -p /path/to/singer/dsasm/vocoder
for frames in 32 384; do
python tools/pack_vocoder_graph_m35.py \
/path/to/voicebank/dsvocoder/nsf_hifigan.onnx \
--frames "$frames" \
--vnni-scope all-k711 \
--residual-scope all3711 \
--out "/path/to/singer/dsasm/vocoder/$frames.dsv35" \
--work "build/packer-$frames"
doneSelecting an extracted provider's lib/libdsasm.so lets OpenUtau discover
bin/dsasm-model-tool from the same package root. Automated deployments may
set OPENUTAU_DSASM_LIBRARY to the library. Source-checkout development may
use DIFFSINGER_ASM_HOME and the explicit OPENUTAU_DSASM_MODEL_TOOL
override.
Published singers use an immutable generation selected by one pointer:
<singer>/dsasm/current.json
<singer>/dsasm/generations/<generation>/bundle.json
<singer>/dsasm/generations/<generation>/acoustic/...
<singer>/dsasm/generations/<generation>/vocoder/32.dsv35
<singer>/dsasm/generations/<generation>/vocoder/384.dsv35
Projects use the fixed renderer IDs DIFFSINGER for ONNX and
DIFFSINGER-ASM for this provider. Selecting ASM reports incompatibility or
runtime unavailability instead of silently changing backend. Switching to ASM
may inspect compatibility silently, but conversion remains an explicit user
action.
ABI v3 rejects configurations it cannot reproduce, including energy conditioning and pitch-controllable vocoders. ASM and ONNX renders use separate WAV cache keys. Both ASM modes declare output compatibility revision 1, so a complete pre-render can satisfy later playback and a completed real-time render can satisfy later batch use. Partial PCM is published only to the active playback session; it enters the complete render cache only after the unique final chunk arrives.
The native engine exposes separate four-worker/32-frame real-time and eight-worker/384-frame batch modes. OpenUtau chooses the mode from whether the render is consumed progressively during playback or as a complete block.
For standalone CLI debugging, text vector files accept whitespace- or comma-separated values; F0 and optional variance curves contain one value per acoustic frame:
build/dsasm-acoustic infer /path/to/singer/dsasm/acoustic \
--tokens tokens.txt \
--durations durations.txt \
--f0 f0.txt \
--language-id 4 \
--speaker-emb /path/to/singer.emb \
--depth 0.6 \
--steps 4 \
--out mel.f32Then synthesize a waveform with a bundle whose fixed frame count exactly matches the CLI input:
build/dsasm-vocoder-m40 infer /path/to/singer/dsasm/vocoder/384.dsv35 \
--mel mel.f32 \
--f0 f0.f32 \
--out wave.f32 \
--wav output.wav \
--workers 8 \
--rounds 3The acoustic CLI does not perform text/phoneme parsing or score preparation.
sum(durations) must equal the F0 frame count, and the vocoder bundle must have
been compiled for that same frame count. A deployed model may also require
language and speaker inputs; inspect reports the enabled features.
dsasm-acoustic accepts F0 as text, while dsasm-vocoder-m40 expects the same
curve as raw native-endian float32 (f0.f32). Mel and waveform .f32 files are
also headerless float32 arrays.
ABI v3 mode configuration records the worker, region, bucket, and overlap
values that passed each architecture's current end-to-end gate. Runtime
defaults select the corresponding shape-gated graph and kernel paths.
config/best-inference.env preserves the common
environment controls used for benchmark reproduction.
. config/best-inference.env
env | grep '^DSASM_' | sortDo not use the promoted AVX-VNNI settings on a CPU without AVX-VNNI. Local
source builds default to -march=native; optional AVX-VNNI vocoder paths
remain runtime-gated.
Build the shared library, validate ABI layout and exported dependencies, and run the real-model streaming test when local packed artifacts are available:
make -j"$(nproc)" engine-check
make model-tool-check
make engine-real-stream-check \
ENGINE_REAL_ACOUSTIC=/path/to/packed/acoustic \
ENGINE_REAL_VOCODER=/path/to/packed/vocoderRun the core native regression checks:
make -j"$(nproc)" m40-1-check m38-check m58-checkRun the promoted residual-kernel benchmark and its parity gate:
make m58-benchThe full real-voicebank E2E comparison needs locally prepared model and fixture artifacts and is documented in docs/PERFORMANCE.md. Voicebanks and singer embeddings are intentionally not distributed here.
The native ABI v3 exposes explicit real-time streaming and block batch modes. Mode configuration is queryable, and mode engines reject mismatched request modes, buckets, or overlap values. See the OpenUtau binding contract.
Exported voicebank (offline)
acoustic.onnx ----> DSFS25 + DSAUX20 + DSLYNX7
nsf_hifigan.onnx -> 32 / 384-frame DSVOC35 buckets
Native inference (online)
OpenUtau score + curves -> stable C ABI -> persistent singer engine
-> FastSpeech2 -> Aux decoder -> Rectified Flow
-> mel + F0 -> bucketed NSF-HiFiGAN
-> PCM chunk callbacks -> MixPlanner -> playback
The packed files are validated and memory-mapped. Large operators use
handwritten assembly; scheduling, graph execution, and control math remain in
small C runtimes. include/dsasm_engine.h is the
stable product boundary; lower-level headers remain implementation-oriented.
- Developer guide
- Performance policy and current evidence
- Dual-architecture optimization strategy
- Real-time streaming performance policy
- Block batch performance policy
- CPU runtime implementation policy
- Stable engine ABI and streaming contract
- Offline model protocol
- OpenUtau supply contract draft
- v0.2.2 release notes
- v0.2.1 release notes
- Manual validation and release process
- Deployment ONNX import
- Real-model acceptance
- Pure-native vocoder executor
- Current residual-fusion milestone
- Packed model formats
The milestone documents preserve the engineering history. New users should start with this README and treat older milestone commands as experiments rather than the recommended release path.
DiffSinger-ASM is an active performance-engineering project from Asmory. The supported surface is intentionally narrow: Linux, x86-64, and exported DiffSinger graphs compatible with the included importers. Model compatibility, quality, and speed should be validated on your own workload before deployment. Validation and releases are run locally; the repository intentionally has no hosted CI or continuous deployment workflow.
No voicebank, singer identity, or third-party model weights are included in the repository or release archives.
Use voice models only with the voice owner's consent. This project follows the upstream DiffSinger responsible-use notice: do not use it to generate a person's voice without permission.
