Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
-
Updated
Aug 23, 2026 - Python
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
Turn detection for full-duplex dialogue communication
A real-time software for turn-taking, backchannel, and head-nodding prediction
[EMNLP 2026 Main] Speaking While Listening: a survey and empirical audit of full-duplex spoken dialogue systems — L0–L3 architectural hierarchy, T×I×R interaction ontology, and a five-state decision machine, with a curated list of models, datasets, and benchmarks.
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
GenPark AI Agent Skill - PSTN/SIP voice telephony state machine managing DTMF tone decoders, call transfer handoffs and IVR navigation.
GenPark AI Agent Skill - Real-time conversational voice activity detection, dynamic silence endpointing and barge-in interruption arbitrator.
GenPark AI Agent Skill - Phonetic Soundex and acoustic confusion matrix resolver correcting domain-specific STT transcription errors.
GenPark AI Agent Skill - Real-time RTP and WebSocket audio packet jitter buffer optimizer smoothing network latency and clock drift for speech streaming.
GenPark AI Agent Skill - Conversational audio acoustic filler injector masking LLM inference and TTS latency with contextual conversational cues.
Python package based on Gibson's framework (2003) for turn-taking in group conversation analysis.
Local-first VAD, barge-in, and turn-taking primitives for interruptible voice agents.
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
Full-duplex conversational turn-taking in C++17 — no ASR, no LLM, no neural VAD. Decides when you have finished speaking from prosody alone, and hears you interrupt it through its own speaker echo.
🏀💬 Advanced User Interfaces course project developed as a 1st year student of MSc in CS Eng. at Politecnico di Milano
Open middleware protocol to make voice AI agents sound human — ambient audio, turn-taking, spoken register. Built on Pipecat.
To associate your repository with the turn-taking topic, visit your repo's landing page and select "manage topics."