FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
-
Updated
Jan 25, 2024 - Python
FunCodec is a research-oriented toolkit for audio quantization and downstream applications, such as text-to-speech synthesis, music generation et.al.
Easily train a good VC model with voice data <= 10 mins!
Pure C inference engine for Qwen3-TTS text-to-speech. No Python, no PyTorch — just C and BLAS. Supports 0.6B and 1.7B models, 9 voices, 10 languages.
a custom comfyui node for coqui-ai/TTS's xtts module! support 17 languages voice cloning and tts
💬 "Realtime" voice transcription and cloning using ElevenLabs's API.
Unity package for using Spark-TTS on-device models. This is a C# port of https://github.com/SparkAudio/Spark-TTS by SparkAudio team and uses converted ONNX models instead of the PyTorch models in the original repo
Using a single image and just 10 seconds of sample audio, our project enables you to create a video where it appears as if you're speaking the desired text.
A fine-tuned SpeechT5 Urdu TTS model with voice cloning that converts both Urdu and Roman Urdu text into natural speech. Trained on diverse Urdu and Zia Mohiuddin recordings, it offers expressive, speaker-specific synthesis with a FastAPI demo for easy testing.
AI-powered audio creation platform offering TTS, Voice Cloning, Voice Changing, dubbing, audiobooks and more.
Clone a voice in 5 seconds to generate arbitrary speech in real-time
Spark-TTS Inference Code and Export to Onnx. The exported onnx models can be used for on-device inference like in https://github.com/genesisinteractive/Spark-TTS-Unity
Audio tour guide website with K-POP celebrity's voice using Text-To-Speech model
Client project for Fontys University of Applied Sciences. Semester 4 Creative Technology
Secure voice AI with biometric authentication & voice cloning. 100+ languages, <200ms latency, hybrid edge computing. "I don't talk to strangers." Apache 2.0
This repository provides a Google Colab notebook for voice cloning using the Coqui XTTS-V2 model. It allows users to clone voices from audio samples and generate speech in multiple languages.
Serverless Voice Cloning (Text-to-Speech) Web App on AWS
Unity package for on-device Qwen3-TTS (12Hz) inference: VoiceDesign and Base ICL voice cloning via ONNX Runtime.
A modular, 100% offline voice cloning audiobook generator built with Python and PyTorch. Features intelligent text-to-speech chunking, local Hugging Face model integration, and automated audio processing for private, high-fidelity PDF narration
Polyglot Echo is an end-to-end, low-latency conversational AI pipeline designed to transcribe user speech, generate context-aware LLM translations, and instantly clone the speaker's voice to read back the response in multiple target languages.
To associate your repository with the voicecloning topic, visit your repo's landing page and select "manage topics."