From b57e07afe940e267a298ecfdd1d03dbf7be9f25a Mon Sep 17 00:00:00 2001 From: Dominic Rubas <1042243+rubas@users.noreply.github.com> Date: Sat, 26 Sep 2026 23:26:47 +0200 Subject: [PATCH 1/2] docs: plain English README and AGENTS.md --- AGENTS.md | 138 ++++++++++++++++++++---------------------------------- README.md | 85 +++++++++++++++------------------ 2 files changed, 88 insertions(+), 135 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 26aff6f..d2ea7ef 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -1,99 +1,63 @@ -# AGENTS.md - -## Goal - -Elixir bindings for whisper.cpp. A Rustler NIF links whisper.cpp through the -`whisper-rs` crate; the Elixir side adds typed options, validation, and PCM -slicing. The Hex package ships precompiled NIF artefacts, so most users never -build Rust. - -Audio decoding stays out of the library. `transcribe/3` takes -`{:pcm_f32, binary}` only (little-endian f32, mono, 16 kHz); a file path or a -bare binary returns `:invalid_request`. Callers decode upstream (ffmpeg, -bumblebee) and share one decoded buffer across a pipeline. - -## Gates - -`ci.yml` runs `task check` on pushes to `main` and on pull requests, so a -push to a branch with no open pull request runs nothing. - -- `task test:integration` downloads `ggml-tiny.en` and the multilingual - `ggml-tiny` (~75 MB each) and runs real inference. `integration.yml` runs it - weekly and on manual dispatch, never per pull request. Run it locally when you - touch the NIF boundary. -- `release.yml` builds the six-entry NIF matrix. One backend already takes - minutes to build from source, so leave the matrix to CI. - -## Layout - -- `checksum-Elixir.WhisperCpp.Native.exs` is tracked on purpose, see Release. - -## House decisions - -- `whisper-rs` and `whisper-rs-sys` resolve through a `[patch.crates-io]` pin to - a vendor branch of this repo, not from crates.io. That branch adds the - callback and CString-leak fixes and moves the whisper.cpp submodule to v1.9.4. - A `whisper-rs` version bump means re-checking the patch, see issue #26. -- One accelerator per build. `WHISPER_CPP_FEATURES` picks the cargo feature and - `WHISPER_CPP_BUILD=1` forces a source build. Precompiled variants exist only - for `cuda` (x86_64 and aarch64 Linux) and `hipblas` (x86_64 Linux); users - select one with `WHISPER_CPP_VARIANT`. The darwin artefact is built with - `metal`. The other features in `Cargo.toml` (`vulkan`, `coreml`, - `intel-sycl`, `openblas`, `openmp`) are source build only. -- A change to `WHISPER_CPP_FEATURES`, `WHISPER_CPP_VARIANT`, or - `WHISPER_CPP_BUILD` recompiles the NIF wrapper. The `build:*` tasks set them - for their own run only, so a later plain `task test` goes back to the - precompiled artefact. Test a backend build with the same variables: +# whisper_cpp + +The README describes the library, its backends, and the CPU baseline. The Hex package ships precompiled NIFs, so most +users never build Rust. + +## Checks + +- `ci.yml` runs `task check` on pushes to `main` and on pull requests. A push to a branch without a pull request runs + nothing. +- `task test:integration` runs real inference. `integration.yml` runs it weekly and on manual dispatch, never on a pull + request. Run it locally when you change the NIF boundary. +- `release.yml` builds six NIFs (target and variant). One backend takes minutes to build from source, so leave the + matrix to CI. + +## Rules + +- Audio decoding stays out of the library. `transcribe/3` takes only `{:pcm_f32, binary}`; a file path or a bare binary + returns `:invalid_request`. Callers decode first and share one decoded buffer across a pipeline. +- No silent fallback. An unknown option, a device the NIF does not have, and an audio shape the library rejects all + return an error, never a default. +- Errors cross the NIF boundary as `{:error, %{type, message, details}}`. `errors.rs` sets the type, and + `WhisperCpp.Error` maps it to a reason atom. An unknown type becomes `:native_error`, so a new type needs both sides. +- Credo enables `Readability.Specs`, so a public function without a `@spec` fails. `Readability.ModuleDoc` is off, but + every public module still gets a `@moduledoc`; no check catches a missing one. +- `whisper-rs` and `whisper-rs-sys` come from a `[patch.crates-io]` pin to the `vendor/whisper-rs-0.16.0-patched` + branch of this repo. It adds the callback and CString-leak fixes and moves whisper.cpp to v1.9.4. On a `whisper-rs` + version bump, check the patch again (issue #26). +- `WHISPER_CPP_FEATURES` picks the cargo feature, and `WHISPER_CPP_BUILD=1` forces a source build. A change to either, + or to `WHISPER_CPP_VARIANT`, recompiles the NIF wrapper. The `build:*` tasks set them only for their own run, so a + later `task test` uses the precompiled NIF again. Test a backend with the same variables, for example `WHISPER_CPP_BUILD=1 WHISPER_CPP_FEATURES=cuda task test:integration`. -- Errors cross the boundary as `{:error, %{type, message, details}}`. - `errors.rs` sets the type and `WhisperCpp.Error` maps it to a reason atom; an - unrecognised type becomes `:native_error`. A new type needs both sides. -- No silent fallback. An unknown option, a device the artefact does not carry, - and an audio shape the library rejects all return an error, never a default. -- Credo runs `--strict` with the ExDNA check and every check in - `ExSlop.recommended_checks/0`, so an ExSlop bump turns on its new recommended - checks. `Readability.Specs` is on, so a public function without a `@spec` - fails the gate. - `Readability.ModuleDoc` is off, so a `@moduledoc` on every public module is a - house rule that no check catches. Write it anyway. ## Pitfalls -- `release.yml` builds with `GGML_NATIVE=OFF` and `GGML_CPU_ARM_ARCH` from the - matrix. Drop them and ggml builds for the runner CPU, so the artefact can die - with SIGILL on an older CPU. The release job fails when the ggml CPU flags - contain `native`. Local source builds keep the native tuning. -- `release.yml` parses `nif_versions:` out of `lib/whisper_cpp/native.ex` with - `sed`. Reformat that line and the release job fails. -- A new precompiled variant needs an entry in both the `@variants` map of - `native.ex` and the `release.yml` build matrix, and each miss breaks - differently. Without the `native.ex` entry the compile fails, because the - variant is not published for the target. Without the matrix entry the - suffix is there but no such tarball was published, so the install fails on - the download. -- The ROCm build needs the `GPU_TARGETS` arch list in `release.yml`. Without - gfx1200 and gfx1201 the artefact loads and reports the GPU, then dies on the - first kernel launch. The release job checks the `.hip_fatbin` size to catch - this before publishing. +- `release.yml` builds with `GGML_NATIVE=OFF` and the `GGML_CPU_ARM_ARCH` of the matrix. Without them ggml builds for + the runner CPU, and the NIF can die with SIGILL on an older CPU. The release job fails when `CMakeCache.txt` does not + have `GGML_NATIVE:BOOL=OFF`. Local source builds keep the native tuning. +- `release.yml` reads `nif_versions:` from `lib/whisper_cpp/native.ex` with `sed`. If you reformat that line, the + release job fails. +- A new precompiled variant needs an entry in the `@variants` map of `native.ex` and in the `release.yml` build + matrix. Without the `native.ex` entry, the compile fails because the variant is not published for the target. + Without the matrix entry, the install fails on the download because no tarball exists. +- The ROCm build needs gfx1200 and gfx1201 in the `GPU_TARGETS` list in `release.yml`. Without them the NIF loads and + reports the GPU, then dies on the first kernel launch. The release job checks the size of the `.hip_fatbin` section + to catch this before it publishes. ## Release -1. Bump `@version` in `mix.exs`, add the `CHANGELOG.md` entry, push to `main`. -2. On every push to `main`, `release.yml` releases when no tag exists for the - `mix.exs` version. It builds a tarball per target and variant, creates the - tag, and uploads the tarballs plus `SHA256SUMS`. A run that is dropped - before it creates the tag does not lose the release, because the next push - retries it. Once the tag exists, a manual dispatch builds the tag's commit - and fails when the tag is missing or its `mix.exs` version differs. The - dispatch uploads only the assets the release does not have yet. It never - replaces a published tarball, because a new tarball breaks the checksum file - in the Hex package. -3. Regenerate the checksum file from the published assets, then commit and push - it. The checksum for each tag stays reproducible from the repo: +1. Bump `@version` in `mix.exs`, add the `CHANGELOG.md` entry, and push to `main`. +2. On each push to `main`, `release.yml` releases when no tag exists for the `mix.exs` version. It builds a tarball per + target and variant, creates the tag, and uploads the tarballs and `SHA256SUMS`. If a run stops before it creates the + tag, the next push retries the release. +3. After the tag exists, a manual dispatch builds the tag's commit. It fails when the tag is missing or its `mix.exs` + version differs, and it uploads only the assets the release does not have yet. It never replaces a published + tarball, because a new tarball breaks the checksum file in the Hex package. +4. Regenerate `checksum-Elixir.WhisperCpp.Native.exs` from the published assets, then commit and push it. This keeps + the checksum of each tag reproducible from the repo: ```bash mix rustler_precompiled.download WhisperCpp.Native --all --no-config --ignore-unavailable --print ``` -4. Run `mix hex.publish` from a clean tree. `mix.exs` ships `checksum-*.exs` - inside the Hex tarball. +5. Run `mix hex.publish` from a clean tree. `mix.exs` puts `checksum-*.exs` in the Hex tarball. diff --git a/README.md b/README.md index da39203..4e6c8c7 100644 --- a/README.md +++ b/README.md @@ -1,10 +1,8 @@ # whisper_cpp -A thin Elixir wrapper around [`whisper-rs`](https://codeberg.org/tazz4843/whisper-rs), -the Rust bindings to [whisper.cpp](https://github.com/ggerganov/whisper.cpp). -It exposes whisper.cpp speech-to-text to the BEAM through a Rustler NIF: load a -model, hand it 16 kHz mono f32 PCM, get structured segments back. No subprocess, -no Python, no temporary files. +Elixir bindings for [whisper.cpp](https://github.com/ggerganov/whisper.cpp) speech-to-text. A Rustler NIF on the +[`whisper-rs`](https://codeberg.org/tazz4843/whisper-rs) crate runs whisper.cpp in the BEAM process. You load a model, +pass it 16 kHz mono f32 PCM, and get back structured segments. No subprocess, no Python, no temporary files. ## Installation @@ -14,15 +12,15 @@ def deps do end ``` -Installation downloads a precompiled NIF for your target from the project's -GitHub releases - no Rust toolchain needed. Requires Elixir 1.19+. +Mix downloads a precompiled NIF for your target from the GitHub releases, so you need no Rust toolchain. The package +needs Elixir 1.19 or newer. ## Usage ```elixir {:ok, model} = WhisperCpp.load_model("models/ggml-large-v3.bin") -# Decode upstream (ffmpeg, bumblebee, ...) into 16 kHz mono f32 PCM: +# Decode the audio first (ffmpeg, bumblebee, ...) into 16 kHz mono f32 PCM: # ffmpeg -i jfk.wav -f f32le -ac 1 -ar 16000 jfk.pcm pcm = File.read!("jfk.pcm") @@ -33,33 +31,33 @@ IO.puts(text) for s <- segs, do: IO.puts("[#{s.start}-#{s.end}] #{s.text}") ``` -Audio is always `{:pcm_f32, binary}` - little-endian f32 samples, mono, 16 kHz, -normalised to `[-1.0, 1.0]`. The library does **not** decode WAV/MP3/etc; -decode upstream. `transcribe_slice/4` runs a `[start_s, end_s)` window of a -master PCM buffer and shifts the returned times back into the source timeline. +Audio is always `{:pcm_f32, binary}`: little-endian f32 samples, mono, 16 kHz, in the range `[-1.0, 1.0]`. The library +does not decode WAV, MP3, or other formats; decode them before the call. `transcribe_slice/4` transcribes a +`[start_s, end_s)` window of a larger PCM buffer and returns times on the timeline of the full buffer. -Built-in silero voice activity detection strips silence before the encoder: -pass `vad_model_path:` (`ggml-silero-v5.1.2.bin`, ~0.85 MB, from -[ggml-org/whisper-vad](https://huggingface.co/ggml-org/whisper-vad)) and -timestamps stay on the original timeline. +The built-in silero voice activity detection removes silence before the encoder. Pass `vad_model_path:` with +`ggml-silero-v5.1.2.bin` (about 0.85 MB, from [ggml-org/whisper-vad](https://huggingface.co/ggml-org/whisper-vad)). +The timestamps stay on the original timeline. -See [the docs](https://hexdocs.pm/whisper_cpp) for the full option list -(`:translate`, `:initial_prompt`, `:word_timestamps`, `:beam_size`, -`:n_threads`, VAD tuning, cancellation, progress messages, ...) and error -handling. +[The docs](https://hexdocs.pm/whisper_cpp) list all options (`:translate`, `:initial_prompt`, `:word_timestamps`, +`:beam_size`, `:n_threads`, VAD tuning, cancellation, progress messages, and more) and the errors. ## Backends -CPU is available in every build except `coreml` (see below). Pick one -accelerator per build; the precompiled Hex package ships CPU plus `cuda` / -`hipblas` variants for Linux and Metal on Apple Silicon, selected via -`WHISPER_CPP_VARIANT`: +Each build has one accelerator. Every build also runs on the CPU, except `coreml`. The precompiled package has a CPU +build for each target, `cuda` and `hipblas` variants for Linux, and Metal on Apple Silicon. `WHISPER_CPP_VARIANT` +selects a variant: ```bash WHISPER_CPP_VARIANT=cuda mix deps.compile whisper_cpp ``` -The precompiled NIFs need this CPU baseline: +| Variant | Targets | +| --------- | -------------------------------------------------------- | +| `cuda` | `x86_64-unknown-linux-gnu`, `aarch64-unknown-linux-gnu` | +| `hipblas` | `x86_64-unknown-linux-gnu` | + +The precompiled NIFs, the GPU variants included, need this CPU: | Target | Minimum CPU | | --------------------------- | ------------------------------------------------------------------- | @@ -67,38 +65,29 @@ The precompiled NIFs need this CPU baseline: | `aarch64-unknown-linux-gnu` | ARMv8.2-A with dotprod and fp16 (Neoverse N1, Cortex-A76, or newer) | | `aarch64-apple-darwin` | Apple M1 or newer | -The `cuda` and `hipblas` variants need the same CPU. On an older CPU, build -from source with `WHISPER_CPP_BUILD=1`. A source build tunes ggml for the CPU -it runs on. - -This baseline applies from 0.5.0 on. Releases up to 0.4.1 were tuned -for the CPU of the release runner. Their `aarch64-unknown-linux-gnu` artefacts -need SVE and i8mm, and the 0.3.1 `x86_64-unknown-linux-gnu` artefact needs -AVX-512. On a CPU without these, build those versions from source. +On an older CPU, build from source with `WHISPER_CPP_BUILD=1`. A source build tunes ggml for the CPU it runs on. This +baseline applies from 0.5.0. Up to 0.4.1, the `aarch64-unknown-linux-gnu` NIFs need SVE and i8mm, and the 0.3.1 +`x86_64-unknown-linux-gnu` NIF needs AVX-512. -To build from source with any whisper-rs backend (`cuda`, `hipblas`, `vulkan`, -`metal`, `coreml`, `intel-sycl`, `openblas`, `openmp`): +A source build can use any `whisper-rs` backend: `cuda`, `hipblas`, `vulkan`, `metal`, `coreml`, `intel-sycl`, +`openblas`, or `openmp`. ```bash WHISPER_CPP_BUILD=1 WHISPER_CPP_FEATURES=cuda mix deps.compile whisper_cpp ``` -Source builds need Rust 1.98 or later, `cmake`, a C++17 compiler, and the -backend's own SDK (CUDA toolkit, ROCm, Vulkan SDK, ...). +A source build needs Rust 1.98 or newer, `cmake`, a C++17 compiler, and the SDK of the backend (CUDA toolkit, ROCm, +Vulkan SDK, ...). -A `coreml` build uses the Core ML encoder whenever the model's -`-encoder.mlmodelc` is present and cannot turn it off per model. It rejects -`device: :cpu` and `use_gpu: false` with `:invalid_request`; build without -`coreml` for CPU-only inference. +A `coreml` build uses the Core ML encoder whenever the model's `-encoder.mlmodelc` exists, and you cannot turn it off +per model. It returns `:invalid_request` for `device: :cpu` and `use_gpu: false`. For CPU-only inference, build without +`coreml`. -## Testing +## Development -```bash -mix test # unit tests, no downloads -mix test --include integration # downloads ggml-tiny.en + ggml-tiny, real inference -``` +`task check` runs the format check, compile, lint, the Elixir and Rust unit tests, and `zizmor` on the workflows. +`task test:integration` downloads `ggml-tiny.en` and `ggml-tiny` (about 75 MB each) and runs real inference. ## License -MIT. whisper.cpp is MIT-licensed; `whisper-rs` is public domain (Unlicense) -and vendors whisper.cpp, linking it statically. +MIT. whisper.cpp is MIT. `whisper-rs` is public domain (Unlicense); it vendors whisper.cpp and links it statically. From b4767ef36d05b7a5104a7377d73bfe21f44674a8 Mon Sep 17 00:00:00 2001 From: Dominic Rubas <1042243+rubas@users.noreply.github.com> Date: Sat, 26 Sep 2026 23:27:06 +0200 Subject: [PATCH 2/2] docs: align the variant table --- README.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 4e6c8c7..73136ca 100644 --- a/README.md +++ b/README.md @@ -52,10 +52,10 @@ selects a variant: WHISPER_CPP_VARIANT=cuda mix deps.compile whisper_cpp ``` -| Variant | Targets | -| --------- | -------------------------------------------------------- | +| Variant | Targets | +| --------- | ------------------------------------------------------- | | `cuda` | `x86_64-unknown-linux-gnu`, `aarch64-unknown-linux-gnu` | -| `hipblas` | `x86_64-unknown-linux-gnu` | +| `hipblas` | `x86_64-unknown-linux-gnu` | The precompiled NIFs, the GPU variants included, need this CPU: