Voice agent infrastructure
Facts from public sources. Notes from production.
A curated directory of voice models, installable skills, and agent infrastructure — structured for builders, not marketing decks.
STT, TTS, and STS models plus audio tools, voice AI, and related models from 27 labs.
163
converts spoken audio into text. Sometimes called "ASR" (automatic speech recognition). Deepgram, AssemblyAI.40converts written text into spoken audio. ElevenLabs, Cartesia.94converts spoken audio directly into spoken audio, skipping the intermediate text step. OpenAI Realtime, Ultravox.16software that determines when someone is speaking vs. silent. Critical for knowing when to interrupt or wait.4Noise cancellation4
Installable agent skills for AI coding tools, focused on voice workflows and integrations.
Top TTS models
Ranked by TTS Arena and Artificial Analysis benchmarks where available, then latency.
Order by
| Model | Lab | Type | Benchmark |
|---|---|---|---|
| Cartesia | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| Rime | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| Rime | converts written text into spoken audio. ElevenLabs, Cartesia. | AA indexed | |
| Rime | converts written text into spoken audio. ElevenLabs, Cartesia. | AA indexed | |
| KugelAudio | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| KugelAudio | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| KugelAudio | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| Murf | converts written text into spoken audio. ElevenLabs, Cartesia. | AA indexed | |
| PlayHT | converts written text into spoken audio. ElevenLabs, Cartesia. | — | |
| AWS | converts written text into spoken audio. ElevenLabs, Cartesia. | AA indexed |
Browse by type
Jump to the models directory with a type filter applied.
converts spoken audio into text. Sometimes called "ASR" (automatic speech recognition). Deepgram, AssemblyAI.40converts written text into spoken audio. ElevenLabs, Cartesia.94converts spoken audio directly into spoken audio, skipping the intermediate text step. OpenAI Realtime, Ultravox.16software that determines when someone is speaking vs. silent. Critical for knowing when to interrupt or wait.4Noise cancellation4Voice isolation2standalone products that detect and mask or tokenize sensitive content (PII) in transcripts, LLM payloads, or audio-adjacent streams—not the same as an STT vendor’s “PII redaction” checkbox on transcription output.2converts speech in one language to text or speech in another, often in real-time.1