Speech and realtime AI
Voice as a product surface: transcription, synthesis, and agents a person can hold a conversation with. Written for roles naming “AI solutions for communication, audio/video streaming, chatbots”.
| # | File | Covers |
|---|---|---|
| 01 | Speech: STT and TTS | streaming vs batch ASR, WER and why it misleads, diarization, custom vocabulary, streaming TTS, cost per minute, testing |
| 02 | Realtime voice agents | cascaded vs speech-to-speech, the latency budget, turn detection and barge-in, WebRTC transport, Pipecat/LiveKit, tools and escalation, evaluation |
| 03 | Voice platforms and cloning | ElevenLabs, Deepgram, AssemblyAI; self-host vs buy; consent, disclosure, and voice as biometric data |