End-to-end documentation to set up your own local & fully private LLM server on Debian. Equipped with chat, web search, RAG, model management, MCP servers, image generation, and TTS.
-
Updated
Jun 30, 2026
End-to-end documentation to set up your own local & fully private LLM server on Debian. Equipped with chat, web search, RAG, model management, MCP servers, image generation, and TTS.
Kokoro-FastAPI为基础,支持hexgrad/Kokoro-82M-v1.1-zh模型,优化中英文混读,中文语音更自然,支持 Docker 自动化化部署,支持 NVIDIA GPU 加速和 CPU 推理两种运行模式
Zero-shot clone-tuner for Kokoro-82M. Generate stock-compatible custom voice packs from 5-seconds of reference audio in less than 1-second
Send text from browser to Kokoro-FastAPI for TTS generation
Text-to-speech CLI tool that uses the Kokoro model for inference. Runs extremely fast locally with or without a GPU. Render smooth speech faster than real-time on most machines. Use Kokoro from CLI or the FastAPI webserver via HTTP requests or directly in the browser. Supports audio playback from the CLI, web interface, or download in many formats.
Launch and optimize llama.cpp servers automatically across Linux, macOS, and Windows using hardware detection and configuration tuning.
An AI-powered daily digest creator that transforms YouTube videos into high-fidelity audio news programs using Gemini summarization and pluggable TTS backends (Gemini & Kokoro). supports CLI, FastAPI, and MCP.
Highlight. Right-click. Listen. A fully local, private text-to-speech browser ecosystem. No cloud, no API keys, no telemetry. Your text never leaves your machine. Powered by a FastAPI backend running XTTS v2 (zero-shot voice cloning from your own .wav/.mp3 samples) or Kokoro-82M (ultra-low latency), and a Manifest V3 Chrome extension.
To associate your repository with the kokoro-fastapi topic, visit your repo's landing page and select "manage topics."