Published pricing and latency for the models and voice engines used to build AI agents — each figure quoted from the vendor's own page and linked to it.
A dated snapshot of what AI model and voice vendors publish about their own products: token pricing and, where the vendor states one, latency. PxlPeak has not run a benchmark and does not present these as measurements. Vendor-published latency is a best-case figure produced on the vendor's own infrastructure.
Key Takeaways
Published list price per 1M tokens. Capability scores are deliberately absent — PxlPeak runs no evaluation, and any “logic” or “coding” score shown without a named benchmark behind it is an invented number.
| Model | Input / 1M | Output / 1M |
|---|---|---|
Claude Fable 5 Anthropic Anthropic's most capable widely released model, GA 2026-06-09. Added 2026-08-05 -- the index had been missing it entirely, so anything treating this file as the complete model canon could not see it. platform.claude.com— source for Claude Fable 5 pricing, opens in a new tab | $10.00 | $50.00 |
Claude Opus 5 Anthropic platform.claude.com— source for Claude Opus 5 pricing, opens in a new tab | $5.00 | $25.00 |
Claude Sonnet 5 Anthropic platform.claude.com— source for Claude Sonnet 5 pricing, opens in a new tab | $3.00 | $15.00 |
Claude Haiku 4.5 Anthropic platform.claude.com— source for Claude Haiku 4.5 pricing, opens in a new tab | $1.00 | $5.00 |
GPT-5.6-sol OpenAI Standard tier, short context. Was $5.00 / $30.00 when read on 2026-08-03; long-context prompts are billed higher. developers.openai.com— source for GPT-5.6-sol pricing, opens in a new tab | $4.00 | $20.00 |
GPT-5.6-terra OpenAI developers.openai.com— source for GPT-5.6-terra pricing, opens in a new tab | $2.00 | $12.00 |
GPT-5.6-luna OpenAI developers.openai.com— source for GPT-5.6-luna pricing, opens in a new tab | $0.20 | $1.20 |
Gemini 3.1 Pro Preview Google Prompts at or under 200k tokens; $4.00 / $18.00 above that. ai.google.dev— source for Gemini 3.1 Pro Preview pricing, opens in a new tab | $2.00 | $12.00 |
Gemini 3.5 Flash-Lite Google ai.google.dev— source for Gemini 3.5 Flash-Lite pricing, opens in a new tab | $0.30 | $2.50 |
DeepSeek-V4-flash DeepSeek Off-peak cache-miss input / off-peak output. Peak doubles both ($0.44 / $1.32). Cache-hit input is $0.007 off-peak. Was a flat $0.14 / $0.28 on 2026-08-03 — DeepSeek has since split pricing by peak window and cache state, so no single figure describes it. api-docs.deepseek.com— source for DeepSeek-V4-flash pricing, opens in a new tab | $0.22 | $0.66 |
DeepSeek-V4-pro DeepSeek Off-peak cache-miss input / off-peak output. Peak doubles both ($1.32 / $3.96). Cache-hit input is $0.022 off-peak. Was a flat $0.435 / $0.87 on 2026-08-03; the peak-hour pricing this row previously flagged as announced is now live. api-docs.deepseek.com— source for DeepSeek-V4-pro pricing, opens in a new tab | $0.66 | $1.98 |
The figures measure different things. A text-to-speech model latency covers one component; a platform's end-to-end response time covers speech recognition, model inference, speech synthesis and network round trips together. A component number will always look faster than an end-to-end number, and ranking them against each other would be meaningless.
| Product | What the figure measures | Vendor states |
|---|---|---|
Flash / Turbo TTS ElevenLabs elevenlabs.io— source for Flash / Turbo TTS latency, opens in a new tab | Text-to-speech model latency only | ~75ms |
Multilingual v2/v3 TTS ElevenLabs elevenlabs.io— source for Multilingual v2/v3 TTS latency, opens in a new tab | Text-to-speech model latency only | ~250–300ms |
Scribe v2 Realtime STT ElevenLabs elevenlabs.io— source for Scribe v2 Realtime STT latency, opens in a new tab | Speech-to-text streaming latency only | ~150ms |
Aura-2 TTS Deepgram deepgram.com— source for Aura-2 TTS latency, opens in a new tab | Streaming text-to-speech latency only | sub-200ms |
Voice agent platform Vapi docs.vapi.ai— source for Voice agent platform latency, opens in a new tab | Full end-to-end response time, one conversational turn | sub-600ms |
Retell AI:Publishes per-minute pricing and a cost calculator, but no end-to-end latency figure was found on its pricing page.retellai.com— source for Retell AI pricing, opens in a new tab
Vendor specs are best-case. We can run head-to-head tests on your own data and call volume to find the most cost-efficient combination.
Ready?
Call now and talk to Aria, our AI strategist — or book a free 30-minute assessment.
Aria picks up instantly · 24/7 · Free assessment · 30-day guarantee