Skip to content

Repository files navigation

KAM TTS

A voice-cloned text-to-speech Chrome (MV3) extension paired with a local Python/Flask server. It reads web pages aloud using XTTS-v2, with a Whisper-based quality-learning loop that adapts synthesis parameters per chunk fingerprint over time.

The voice it ships with

My own voice ships in server/voice_samples/, so a fresh clone speaks the moment it starts rather than asking you to record sixteen passages before you know whether you like it. The code is MIT; the recordings are not. Use them to run and evaluate KAM TTS, but not in anything you publish, and not to represent me. VOICE.md has the terms and the fifteen-minute path to replacing them with your own, which is the point of the project.

Nothing else of mine ships. The quality database, the learned settings, the tuning history and the drift baselines are all gitignored, so every install starts with a clean slate and learns your reading rather than inheriting mine. The one exception is pronunciation_store.json, a starter dictionary of technical words (JSON → "Jayson", AWS → "A.W.S"), which has nothing personal in it and saves everyone teaching the same acronyms again.

Status

Version 0.9. Everything described below works and the suites cover it, so it is worth using, but I am calling it 0.9 rather than 1.0 for two honest reasons.

  • The voice it ships with is not finished. The standard passages exist to cover a spread of prosody and I have recorded eight of the sixteen, so the clips do not yet span the range they are meant to. Expect the shipped voice to be weakest on the registers passages 9 to 16 cover, which is sustained low pitch, contrastive stress, equations read aloud and steady enumeration under load. Recording your own sixteen fixes this for your copy.
  • Some of it is proven by tests rather than by use. The non-NVIDIA backends (ROCm, MPS, Intel XPU) are logic-tested and have never run on real hardware, and only the "runaway" kind of hallucination has ever actually been caught, so I don't know whether the other detectors are too conservative or those cases simply don't arise on the pages I read.

I read real pages with this on an RTX 5090 Laptop at RTF 0.452, meaning synthesis runs about twice as fast as playback.

Features

  • Voice cloning from your own recordings — 8–12 short clips, averaged by XTTS-v2 into a personal voice. Multiple named voice profiles, switchable live from the dashboard, each with fully isolated learning.
  • Reads real pages properly — equations (X_{t+1} → "X sub t plus 1"), LaTeX, acronyms, numbers, code tokens; in-page highlight bar tracks the spoken chunk, maths included.
  • A closed learning loop — Whisper listens back to every chunk, scores it, and KAM tunes its own synthesis parameters per sentence fingerprint (length, punctuation, wording, complexity). Thumbs and reports feed the same system; the AI page shows quality trends and whether your reports worked. Everything learned is stored per fingerprint, never by moving your sliders, so your Speech Tuning settings stay yours and learning layers on top of them.
  • Runs on whatever you have — NVIDIA (CUDA), AMD (ROCm), Apple Silicon (Metal), Intel Arc (XPU) or plain CPU. KAM detects which backends are present, picks the best one, proves it can actually run before committing to it, and falls back cleanly if it can't.
  • Hardware-adaptive — measures real synthesis speed on first boot, then adapts buffering, standby and listen-back to suit anything from a CPU-only laptop to a high-end GPU. The measurement streams to the dashboard console while it runs.
  • Guards its own output — every synthesised chunk gets checked for the ways XTTS fails, meaning runaway babble, a looped fragment or a sentence cut off mid-word, and is re-synthesised at a steadier temperature rather than played.
  • Local and private — everything runs on your machine; per-install API token, localhost-only server, nothing leaves your computer.

Licensing / commercial use. KAM TTS builds on XTTS-v2. Verify the licence of the exact model weights you use before any commercial use — some XTTS-v2 releases are non-commercial. Whisper, spaCy, PyTorch and Flask are permissively licensed. Voice cloning also carries consent obligations: only clone voices you have the right to use.


Hardware

KAM runs on whatever compute your machine has. On startup it detects every backend PyTorch can see, picks the best one, and then verifies it with a real operation before loading the model. I added that last step because a backend can report itself as available and still fail on the first kernel, which both MPS and fresh ROCm installs do fairly often, so it drops back to CPU immediately rather than breaking halfway through a page.

Backend Hardware Notes
CUDA NVIDIA Fastest. TF32 is enabled automatically on Ampere and newer.
ROCm AMD, Linux PyTorch reports it as "cuda", so KAM tells them apart and skips the NVIDIA-only tuning.
MPS Apple Silicon fp32 with a CPU fallback for the operations Metal is missing. Usable, but slower than CUDA.
XPU Intel Arc Needs the Intel XPU build of PyTorch.
CPU anything Always works. The thread count is capped so listen-back doesn't starve synthesis.

You can force a choice with the KAM_DEVICE environment variable (cpu, cuda, mps, xpu), which is useful if you want to keep a GPU free or work around a driver that's misbehaving.

Speed

The server measures itself. The first time it boots on a given machine it times the real synthesis path, prints the results to the dashboard console as they happen, and adapts its buffering to them. You don't have to run anything.

To measure it by hand, or before installing the extension:

python hardware_profile.py

Both report the real-time factor (RTF), which is seconds of compute per second of audio. They share the same code so the numbers agree.

RTF Meaning
< 0.5 Synthesis comfortably outpaces playback
0.5–1.0 Keeps ahead, fine for continuous reading
1.0–2.0 Slower than playback, so brief buffering
> 2.0 Reading stalls often, a GPU is strongly recommended

Useful flags: --no-test reports the machine without timing anything, and --device cpu measures a specific backend. Set KAM_NO_BENCHMARK=1 if you don't want the server measuring itself on first boot.

Minimum: 4-core CPU, 8 GB RAM. CPU-only inference works but is several times slower than playback, so it's fine for short passages and not for continuous reading. KAM detects this and buffers more deeply, and samples listen-back rather than running it on every chunk, so the two don't fight over the same cores.

Recommended: NVIDIA GPU with 6 GB or more of VRAM and CUDA, plus 16 GB RAM. That gives real-time or faster synthesis.

Under 6 GB of VRAM works but leaves little headroom, so use a shorter idle standby timeout (Speech Tuning, then Idle standby) and the model will release VRAM when it isn't being used.


Setup

One command (recommended). From the server/ folder, using the Python you intend to run KAM with:

python setup_kam.py

This checks your hardware before downloading anything, installs dependencies, fetches the spaCy model, registers the Chrome native-messaging host, checks for voice reference audio, and offers to profile synthesis speed. It is idempotent — safe to re-run.

Nothing to configure by hand: manifest.json pins a public key, so Chrome gives the extension the same id on every machine and the server already knows it.

You must still install PyTorch for your hardware first — it is the one dependency that must match your machine. Get the correct command from pytorch.org: the CUDA build for NVIDIA, the ROCm build for AMD on Linux, the default macOS wheel for Apple Silicon, the XPU build for Intel Arc, or the CPU build for anything else. requirements.txt deliberately leaves torch unpinned for this reason. KAM adapts to whichever you installed.

Manual setup (if you prefer to run the steps yourself)
  1. Install PyTorch for your hardware from pytorch.org.

  2. Install the rest:

    pip install -r requirements.txt
    python -m spacy download en_core_web_sm
    
  3. Add voice reference clips. Easiest from the dashboard: click ● Record in the top bar and read the passages aloud. It records, trims, checks and saves each clip for you, so there is nothing to install and no files to move. See Recording a voice below.

    If you would rather do it yourself, drop WAV files of 8–15 seconds each into server/voice_samples/ and check them with python check_voice_clips.py. XTTS averages them into a single speaker embedding — separate clips clone better than one long file. A single server/my_voice.wav also works.

  4. Register the native-messaging host (no arguments needed):

    python register_host.py
    

    Only the dashboard power button depends on this. Without it everything else still works, you just start the server yourself with python server.py.

    You normally never need this: the server checks its own registration on every boot and repairs it, so installing once and later moving the project is fine. Start the server and it fixes itself.

    What it installs, on Windows, is a small kam_host.exe into %LOCALAPPDATA%\KAMTTS — deliberately outside the project, since the registry stores absolute paths and those do not follow a folder you move. The launcher has nothing about your machine compiled into it; it reads kam_host.cfg beside it at run time, so a move rewrites two lines rather than rebuilding anything. Each new build is sent a real message and has to answer before it replaces the installed one, and a server that finds an older launcher installed rebuilds it on boot, so fixes to the launcher reach existing installs without anyone re-running this.

Then load the extension: chrome://extensions → Developer mode → Load unpacked → select the extension/ folder. Start the server from the dashboard power button, or python server.py.

About the extension id. Chrome normally derives it by hashing the folder it loaded the extension from, which would give every user a different id and mean the server could not know in advance whose requests to trust. So manifest.json pins a public key and Chrome derives the id from that instead: everyone gets mdhbimlofbadmgombcdmnmnebgglalob, and CORS works with no setup. If you fork this and replace the key, either run python register_host.py <YOUR_ID> or set KAM_EXTENSION_ID, or the server will refuse your build.


Updates

The circular arrow at the right of the dashboard's top bar checks GitHub for a newer version, and the dashboard also checks by itself once, a few seconds after it opens, and never while it stays open. Opened again within six hours of a check, it does not ask at all. When there is one, the arrow becomes a gold Update to x.y.z button and a bar under the top bar says what changed, with What's new, Not now and Update now.

Update now asks the server to install it, then restarts the server and the extension, and the dashboard comes back saying what it was updated from. What installing means depends on how you got KAM TTS:

  • A release zip. It downloads the new release, refuses it unless its SHA-256 matches the one GitHub publishes, and copies it over your copy. Your pronunciations, corrections and voice recordings are never replaced, and if anything fails part-way every file already replaced is put back.
  • A git clone. It fast-forwards to the branch you track. If that would clash with your own edits or commits it stops, says which, and leaves the checkout exactly as it was.

Python packages an update needs are installed first, before any file changes, with torch and torchaudio held at the versions you installed for your hardware, so an update never swaps in a torch build that does not match your machine. Not now hides the bar for that version; the gold button stays.

Recording good reference clips

Clone quality is set almost entirely by the reference audio, so this matters more than any parameter:

  • The 16 standard passages (12–20 clips is the sweet spot, dashboard → ● Record) cover declaratives, questions, numbers, long clauses, lists, warmth, explanation, closings, exclamations, quoted speech, asides and fragments — the full prosodic range KAM reads daily.
  • 8–15 seconds per clip. More clips is not automatically better — the encoder averages them, so one poor clip drags the result down.
  • Quiet room, consistent mic distance (~15–20 cm), no fans or traffic.
  • No processing — no noise gates, compression, EQ or "enhancement".
  • Don't clip. Slightly quiet beats peaking.
  • Read the way you want KAM to read to you. The clone mirrors your delivery, not just your timbre.
  • Vary structure: statements, a question, a list, a longer sentence with clauses, and something with numbers or an acronym.

Optionally add voice_samples/transcripts.txt (one line per clip). On startup the server analyses their prosody and logs which structures your reference audio actually covers.

Clips are screened automatically

XTTS averages every clip in a folder into one speaker embedding, so a single clipped, near-silent or noisy recording drags the whole voice down and there's no way to hear which one did it.

So KAM measures each clip before computing latents and excludes the unusable ones, naming them and saying why in the console. Since the clips are averaged, dropping a bad one can only improve things. It never rejects all of them, and if nothing passes it uses everything and prints the reasons instead.

You can check any profile without switching to it:

python check_voice_clips.py voices/my-voice

Same measurements and same thresholds as the server, so this tells you in advance exactly what the server is going to do.

Speech quality

Two things keep quality even across voice profiles, rather than it depending on how carefully a given voice happened to be recorded:

  • The clip gate above, so a profile recorded later on a laptop mic is held to the same standard as the first one.
  • Output validation. XTTS is autoregressive and sometimes fails to emit a stop token. It doesn't raise an error when that happens, it just returns audio, and the audio is babble or a repeated fragment or a sentence cut off mid-word. So KAM checks the waveform against the text it was meant to speak, and if it fails it re-synthesises at a lower temperature and a higher repetition penalty instead of playing it. Two retries, since sampling failures are random and a colder re-roll almost always works.

GET /quality/rejections reports how many were caught and how they failed, which gives you the hallucination rate as a number rather than an impression. GET /diagnose/<chunk_id> explains any single chunk: how it was labelled, which parameters were used and where they came from, what it cost, how it scored, and which learned rules rewrote it.

Rejected attempts are also kept as evidence. A chunk that fails at temperature 0.45 and then succeeds at 0.29 is a controlled comparison, since the text is identical, one variable changed, and I know the outcome on both sides. It's the only causal data the system produces. Ordinary reading never moves the sampling parameters, so without it the self-tuner has almost nothing to work out a direction from.

How learning is stored

Two rules keep the loop honest and reversible:

  • Learning never writes your settings. Everything KAM works out lives in good_settings.json, keyed by chunk fingerprint, and gets applied on top of your live values at synthesis time. The Speech Tuning sliders stay yours.
  • Reinforcement uses what the chunk was actually made with. A thumbs-up reinforces the parameters stored on that chunk's row rather than whatever the sliders read when you clicked, because a read spans many chunks and those two are often different.

You can see both per chunk through /diagnose/<chunk_id>, which reports where each parameter came from: either a value learned for that fingerprint, or your own defaults.


Recording a voice

Click ● Record in the dashboard top bar. Everything happens in that one screen, so no separate recording software and no moving files around.

  • Pick the profile you are recording into, or make a new one, from the dropdown at the top.
  • Pick a passage from the list on the left, or choose Write my own passage and type anything you like. The text is editable either way, and whatever is on screen is saved next to the clip as a .txt, so every recording has a known transcript. That is worth having: the punctuation says where you paused, and a clip with known text can be checked against what was actually said.
  • Record, watching the level meter. Green is healthy, amber is close to the ceiling, red is clipping, which is the one fault you cannot hear yourself.
  • Trim the take. Almost every recording opens with a breath and ends with the click of the mouse going back to the stop button, and both otherwise end up in the clip that gets cloned. Drag the handles, or press Auto-trim, which looks for where the level stays up rather than where it is merely loud, so a short click at the end is not mistaken for speech. ▶ Play previews exactly what will be saved.
  • Save, and the server measures the clip and says straight away whether it is usable, naming the problem if not. A dot next to each passage tracks what is done.
  • ↓ WAV downloads any clip, since the files are still the real artefact and you should never be locked out of them.

Six good clips make a solid voice, and 12 to 20 is the sweet spot. Press Use this voice when you are done, which rebuilds the speaker embedding without restarting the server.

Scanning a page with your phone

Press 📷 Scan in the dashboard and point your phone's camera at the code. The phone gets a page for taking photos; each one comes back here, gets straightened and read, and the text lands in Custom Text in the popup where you can edit it before it speaks.

Tell it how many columns the page has before you scan. Tesseract can work columns out for itself until it gets it wrong, and when it does it interleaves them line by line into something that reads smoothly and means nothing, so the setting is explicit rather than guessed.

It also strips what you would not want read aloud: page numbers, running heads, the superscript letters that mark footnotes, and verse numbers. Verse numbers are the awkward one, because "40" in a verse marker and "40" in a sentence are the same two characters. They are told apart by counting: markers run in sequence down a page and a quantity does not, so a number is only dropped when it is the one the sequence expects.

How the phone reaches the computer. The server normally listens on 127.0.0.1 and nothing else, which is why a phone cannot see it. Opening a scan session starts a second listener on this machine's wifi address, and closing the panel stops it. That listener serves its own two routes and nothing more, so the rest of the API is not on the network even for as long as a session is open, and the phone is given a random per-session key rather than your API token. Sessions close themselves after twenty minutes. The photos are deleted once their text has been read; nothing is uploaded anywhere and the recogniser runs in the browser.

Both devices need to be on the same wifi. If the computer is on a wired network the phone cannot reach it and the panel will say so rather than showing a code that will not work.

Voice profiles

Switching is instant once a voice's latents are cached, and each voice learns independently — tuning, quality history and baselines never bleed between voices. Only word pronunciations are shared. Clips recorded outside the dashboard can be cleaned with python clean_voice_clips.py --in <raw> --out <folder>.

Rename a voice with ✎ and remove one with 🗑, both in the voice menu. A rename carries everything with it, since a profile's name is the key on its clip folder, its cached latents, its drift baseline, its rows in the database and its learned settings, and moving only the folder would leave a voice that looks the same and has quietly forgotten what it knew. Deleting throws away recordings that cannot be made again from here, so it tells you how many clips and how many observations are about to go and asks first. The default voice cannot be renamed or deleted because it is the base voice_samples/ folder, and the voice currently being spoken cannot be deleted until you switch away from it.


Security

  • The local API token is generated per install on first server run and stored in server/kam_token.txt (gitignored). The extension fetches it from the server's /token endpoint — no shared secret ships in source.
  • The extension ID is pinned by a public key in manifest.json, so it is the same on every machine and the server knows which origin to trust without any setup. Only the public half is in the repo, which is how every Chrome Web Store extension works; the private key signs .crx builds and is not needed to run KAM.
  • The server binds to 127.0.0.1 only; CORS restricts callers to the extension origin. The one exception is a scan session, which starts a second listener on this machine's wifi address so a phone can reach it. That listener exists only while the Scan panel has a session open, serves two routes and nothing else, is a separate app from the main API rather than the same one behind a filter, and takes a random per-session key that is not the API token. It closes itself after twenty minutes. See Scanning a page with your phone above.
  • Override with the KAM_TOKEN or KAM_EXTENSION_ID environment variables for custom setups.
  • The only request KAM TTS makes on its own account is the update check: the dashboard asks GitHub's releases API which version is newest, sending nothing about you, your voice or what you read. A download happens only after Update now, and is refused unless its SHA-256 matches GitHub's.

Environment variables

All optional — KAM works with none of them set.

Variable Effect
KAM_DEVICE Force a backend: cpu, cuda, mps, xpu. Ignored (with a log line) if that backend isn't present.
KAM_CPU_THREADS Override the CPU thread cap used for inference.
KAM_NO_BENCHMARK Set to 1 to skip the first-boot speed measurement.
KAM_TOKEN Use a fixed API token instead of the per-install generated one.
KAM_EXTENSION_ID Override which extension origin is allowed through CORS. Accepts several IDs separated by commas. Only needed if you replaced the pinned key in manifest.json with your own.
KAM_PYTHON Python interpreter the native host launches the server with.
KAM_SERVER_PY Path to server.py if it isn't next to kam_host.py.
KAM_FRESH_LATENTS Set to 1 to recompute voice latents instead of using the cache.

Troubleshooting

The dashboard opens but nothing loads, or every request fails. That is the extension ID not matching. It should not happen, since the ID is pinned by the key in manifest.json, but it will if you edited or removed that key. The server prints which origin it trusts on every boot, so compare the first few console lines against the ID at chrome://extensions. If they differ, run python register_host.py <YOUR_EXTENSION_ID> and restart the server.

The power button says "Specified native messaging host not found". Chrome could not use the launcher it was pointed at. Start the server once by any other means — Start KAM TTS.bat, or python server.py — and it repairs its own registration on boot, which covers the usual causes: the project folder moved, the interpreter changed, or the launcher was built from older code.

Then quit Chrome completely and reopen it. Closing every window is not the same: while "Continue running background apps" is on, Chrome keeps running in the tray, so check the Chrome icon there or chrome://settings/system. A full quit has cleared this error every time it has been seen with a registration that was otherwise correct.

If it still fails, %LOCALAPPDATA%\KAMTTS\host.log records every launch, so it says whether Chrome started the launcher at all, which separates "Chrome would not start it" from "it started and something went wrong afterwards".

Chrome starts the launcher through cmd.exe, so a Windows username containing &, %, ^ or brackets can break the path on the way in. Registration writes the short 8.3 form of such a path, and warns if the drive has short names turned off.

It says "No GPU in use" but I have one. The console prints which backends it found and why one was rejected. Usually it's a PyTorch build that doesn't match the hardware, like a CPU wheel on an NVIDIA machine, so reinstall torch from pytorch.org for your setup.

A GPU was detected and then dropped to CPU. The backend passed detection and then failed a real operation, and the console gives you the error. This is common on Apple Silicon with older torch versions and on partial ROCm installs. Use KAM_DEVICE to force it if you think it's wrong.

Reading stalls and buffers. Check the measured RTF on the dashboard. Anything above 1.0 means synthesis is slower than playback on this machine, which is expected on CPU.

A voice sounds worse than the default one. Run python check_voice_clips.py voices/<name>. It's almost always the clips, and the report names the specific problem in each file.

Chunks occasionally sound garbled. GET /quality/rejections shows how many were caught and re-synthesised. If the count is high, lower the temperature in Speech Tuning, since a high sampling temperature is what drives these failures.


What is and is not committed

The one piece of personal data that ships is the default voice in server/voice_samples/, deliberately, on the terms in VOICE.md.

Everything else is gitignored and created per install: the token file, the SQLite quality database, learned settings, drift baselines, cached latents, any voice profiles you add under voices/, logs, and the generated native-host manifest. So you inherit the voice and nothing else — no reading history, no tuning, no rules. Your copy learns you from scratch.

About

Text to speech in a cloned voice, running locally. Chrome extension plus a local XTTS-v2 server.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages