Sourced pricing · verified 2026-09-18
What a minute of voice actually costs
Every realtime speech-to-speech model on Sierra's τ³-Voice leaderboard, plus the newer releases it hasn't rated yet — priced on both axes, with the arithmetic shown and the gaps left as gaps.
different billing units in one market: per audio token, per minute, per hour, per character. No vendor publishes all of them.
gpt-live-1 ranks #1 at 81.7% — and is billed flat per minute. It has no token price at all.
OpenAI's cached audio input ($0.40/1M) against uncached ($32.00/1M). Any single $/min figure is meaningless without stating this.
models sit on the efficient frontier — and Gemini 3.8 Live lands on the cheaper one's exact price, with no τ³-Voice score published to place it.
The finding that shapes this page: You want price per minute and price per 1M tokens. Most vendors publish exactly one of the two. Where a vendor publishes a token price and a documented token-per-second rate, the per-minute figure here is derived and marked as such. Where a vendor bills flat per minute and never discloses a token rate — now true of the top three models on the leaderboard — the per-1M-token cell is not applicable, not estimated. Filling those cells would mean inventing numbers.
Assumptions you can change
A “minute of conversation” is not one thing. These two controls drive every derived figure.
Flat per-minute models don't move when you drag this. Token-billed ones do — output audio is the expensive half.
Applies only where a vendor publishes a cached audio rate. Google publishes none for Live audio, so Gemini rows are unaffected.
Violet marks sit on the efficient frontier — nothing is both cheaper and better. Hollow marks in the right-hand rail score on the benchmark but publish no usable price, so they cannot be placed on a price axis at all.
Priced, but not yet rated on τ³-Voice Overall
These have a real, computable cost per minute and no score on this benchmark's Overall table. They are the mirror image of the hollow marks in the chart's right-hand rail, which have a score and no price.
Realtime speech-to-speech
Ranked by τ³-Voice Overall pass^1. Interaction metrics are the leaderboard's own measurements of conversational behaviour — the closest thing to a reproducible “does it feel human” number.
| # | Model | Billing | $/1M audio in | $/1M audio out | $/min | Pass^1 | Latency | Interrupts | Respons. | |
|---|---|---|---|---|---|---|---|---|---|---|
| #1 | gpt-live-1 backend gpt-6 astra (medium) OpenAI · Sep 10, 2026 | per minute | n/a | n/a | $0.0500 published | 81.7% | 2.54s | 22.0% | 95.5% | |
| #2 | Pine Voice Preview reasoning enabled Pine AI · Aug 10, 2026 | not published | n/a | n/a | n/a | 75.4% | 2.14s | 46.9% | 92.7% | |
| #3 | grok-voice-think-fast-1.0 reasoning enabled xAI · Apr 23, 2026 | per minute | n/a | n/a | $0.0500 published | 67.3% | 1.21s | 19.7% | 100.0% | |
| #4 | grok-voice-think-fast-2.0 high xAI · Jul 29, 2026 | per minute | n/a | n/a | $0.0800 published | 62.5% | 1.71s | 10.7% | 100.0% | |
| #5 | qwen3.5-omni-plus-realtime Qwen · Mar 30, 2026 | not published | n/a | n/a | n/a | 53.7% | 1.70s | 36.1% | 99.2% | |
| #6 | gemini-3.1-flash-live-preview thinking high Google · Mar 26, 2026 | per token | $3.00 | $12.00 | $0.0112 derived | 43.8% | 3.15s | 18.9% | 84.7% | |
| #7 | gpt-realtime-2 xhigh OpenAI · May 7, 2026 | per token | $32.00 cached $0.40 | $64.00 | $0.0480 derived | 42.4% | 1.98s | 20.8% | 95.1% | |
| #8 | gpt-realtime-2 minimal OpenAI · May 7, 2026 | per token | $32.00 cached $0.40 | $64.00 | $0.0480 derived | 38.5% | 1.44s | 18.7% | 99.8% | |
| #9 | grok-voice-fast-1.0 xAI · Dec 17, 2025 | not published | n/a | n/a | n/a | 38.3% | 1.15s | 84.3% | 91.3% | |
| #10 | gpt-realtime-1.5 OpenAI · Feb 23, 2026 | per token | $32.00 cached $0.40 | $64.00 | $0.0480 derived | 35.3% | 1.39s | 13.5% | 99.9% | |
| #11 | Cascaded baseline ASR + LLM + TTS Multiple · May 19, 2026 | not published | n/a | n/a | n/a | 31.2% | 4.24s | 64.4% | 79.1% | |
| #12 | gpt-realtime-1.0 OpenAI · Aug 28, 2025 | per token | $32.00 cached $0.40 | $64.00 | $0.0480 derived | 30.4% | 1.53s | 12.3% | 97.8% | |
| #13 | gemini-3.1-flash-live-preview thinking minimal Google · Mar 26, 2026 | per token | $3.00 | $12.00 | $0.0112 derived | 28.6% | 1.64s | 9.8% | 97.0% | |
| #14 | gemini-live-2.5-flash-native-audio Google · Dec 12, 2025 | per token | $3.00 | $12.00 | $0.0112 derived | 25.8% | 1.43s | 20.6% | 80.6% | |
| — | gemini-3.8-live-extended-thinkingnewest multi-step reasoning; speaks while thinking Google · Sep 15, 2026 Vendor-published, other benchmarks: 82.6 AA quality (#1) · 68.6% τ-Voice agentic · 35.1% Sierra Banking · 97.7% Big Bench Audio | per token | $3.00 | $12.00 | $0.0112 derived | not rated | n/a | n/a | n/a | |
| — | gemini-3.8-livenewest built for scale and cost efficiency Google · Sep 15, 2026 Vendor-published, other benchmarks: 2nd on Speech Agent Arena · 97 languages, switched mid-conversation | per token | $3.00 | $12.00 | $0.0112 derived | not rated | n/a | n/a | n/a | |
| — | gpt-realtime-2.1 not yet on τ³-Voice OpenAI · 2026 | per token | $32.00 cached $0.40 | $64.00 | $0.0480 derived | not rated | n/a | n/a | n/a | |
| — | gpt-realtime-2.1-mini not yet on τ³-Voice OpenAI · 2026 | per token | $10.00 cached $0.30 | $20.00 | $0.0150 derived | not rated | n/a | n/a | n/a | |
| — | gpt-audio not realtime; not on τ³-Voice OpenAI · 2026 | per token | $32.00 | $64.00 | $0.0480 derived | not rated | n/a | n/a | n/a | |
| — | gpt-audio-mini not realtime; not on τ³-Voice OpenAI · 2026 | per token | $10.00 | $20.00 | $0.0150 derived | not rated | n/a | n/a | n/a | |
| — | gemini-2.5-flash-native-audio not yet on τ³-Voice Google · Dec 2025 | per token | $3.00 | $12.00 | $0.0112 derived | not rated | n/a | n/a | n/a |
Reading the interaction columns: Latency is time to respond (lower better). Interrupts is how often the model talks over the user (lower better). Responsiveness is whether it responds at all when it should (higher better). n/a means the vendor publishes nothing on that axis.
Everyone else, on Google's benchmarks
Google cited five benchmarks for its newest model and none of them was τ³-Voice Overall. Here is every model scored on each one Google chose — plus, last, the one it left out.
Bars start at zero, which is why most of them look alike — that is the finding, not a rendering fault. Purple bars are Google's. The last panel is the benchmark Google did not quote.
Model names differ by source and are printed exactly as each source publishes them. Artificial Analysis rates Gemini 3.8 Live Extended Thinking; Sierra's Overall table has no 3.8 entry at all, so Google's best-placed model there is the 3.1 Flash Live predecessor. Those are different models and are not stacked into one bar. Separately, GPT-Live-1 appears as 67.9% on Artificial Analysis and 81.7% on Sierra — same model, different harness. That 13.8-point gap is why nothing here is merged onto a shared axis.
One-way models: speech out, speech in
Kept in a separate table on purpose. These bill per character or per minute of audio and do not hold a conversation, so putting them on the same axis as the models above would be a false comparison.
| Model | Provider | Job | Billed in | Rate | Text in / 1M |
|---|---|---|---|---|---|
| tts-1 | OpenAI | Text to speech | 1M characters | $15.00 | — |
| tts-1-hd | OpenAI | Text to speech | 1M characters | $30.00 | — |
| gpt-4o-mini-tts | OpenAI | Text to speech | 1M audio out tokens | $12.00 | $0.60 |
| gemini-2.5-flash-preview-tts | Text to speech | 1M audio out tokens | $10.00 | $0.50 | |
| gemini-3.1-flash-tts-preview | Text to speech | 1M audio out tokens | $20.00 | $1.00 | |
| gemini-2.5-pro-preview-tts | Text to speech | 1M audio out tokens | $20.00 | $1.00 | |
| gpt-transcribe | OpenAI | Speech to text | minute of audio | $0.0045 | — |
| gpt-4o-mini-transcribe | OpenAI | Speech to text | minute of audio | $0.0030 | $1.25 |
| whisper | OpenAI | Speech to text | minute of audio | $0.0060 | — |
| gpt-4o-transcribe | OpenAI | Speech to text | minute of audio | $0.0060 | $2.50 |
| gpt-4o-transcribe-diarize | OpenAI | Speech to text + speakers | minute of audio | $0.0060 | $2.50 |
| gpt-live-transcribe | OpenAI | Realtime transcription | minute of audio | $0.0170 | — |
| gpt-realtime-whisper | OpenAI | Realtime transcription | minute of audio | $0.0170 | — |
| gpt-realtime-translate | OpenAI | Realtime translation | minute of audio | $0.0340 | — |
How they actually sound
What people report about quality and lifelikeness. That is opinion, and it is kept structurally separate from every number above.
Reported reactions from reviews, vendor posts and developer forums. Every claim carries a numbered source. The measured numbers sit under each entry for contrast.
gpt-live-1
OpenAI · #1 on τ³-Voice at 81.7%
- Reviewers describe it as dropping the "robotic turn-taking" of earlier voice AI — it listens and speaks at once rather than trading turns. [7]
- Language-learning app Speak reported wrong interruptions cut by roughly 80% after switching. [7]
- Scores 97.3% on Artificial Analysis' conversational-dynamics measure, the highest reported. [6]
- A launch-day tally put reaction at 78% positive across 482 responses — critics named latency and occasional wrong answers. [7]
- Its measured 2.54s latency is the second-slowest of any model on the leaderboard, despite the naturalness praise. [3]
Gemini 3.8 Live / Live Extended Thinking
Google · released Sep 15, 2026 · absent from τ³-Voice; predecessor scores 43.8%
- Uses early verbal cues — acknowledging with something like "let me check that" — specifically so it doesn't read as a system waiting to answer. [9]
- Extended Thinking reasons and speaks simultaneously, which is the architectural claim developers responded to rather than a demo trick. [9]
- Google reports #1 on Artificial Analysis' Speech-to-Speech Quality Index at 82.6, and 97.7% on Big Bench Audio. [9]
- Detects and switches between 97 languages mid-conversation — a lifelikeness axis nothing else here claims. [9]
- Every score Google published is on a benchmark other than τ³-Voice Overall, so it cannot be ranked against the models above it. [9]
- Developers flag latency gaps and patchy regional availability, and want speech-specific benchmarks before trusting headline numbers. [9]
- Its rated predecessor is the slowest model on the board at 3.15s with the second-lowest responsiveness — unproven that 3.8 fixes this. [3]
grok-voice-think-fast-2.0
xAI · 62.5% on τ³-Voice
- Tuned toward real conversational habits: shorter sentences, one question at a time, less filler. [4]
- The only top-five model to break one second on Artificial Analysis, cutting latency from 1.25s to 0.70s. [6]
- Users report graceful handling of background noise mid-sentence rather than derailing. [4]
- Priced 60% above Think Fast 1.0 ($0.05 → $0.08/min) while scoring lower than 1.0 on this benchmark. [3][4]
- Selectivity drops to 35.9% from 1.0's 51.5% — it is less discriminating about when to speak. [3]
gpt-realtime family
OpenAI · 25.8–42.4% on τ³-Voice
- Instruction-followable delivery — tone and pacing can be directed in the prompt. [1]
- Deepgram's independent VAQI evaluation gives it the best miss-rate in its test set. [10]
- A developer-forum thread titles the 1.5 release a "major regression in voice expressiveness", reporting accents largely gone. [8]
- Practitioners building outbound voice agents describe a prosody plateau — acceptable for low-stakes assistants, not for callers who did not opt in. [10]
- One engineer's writeup of building on the Realtime API is titled around the voices sounding robotic, and moves to a hybrid stack. [10]
Pine Voice Preview
Pine AI · #2 on τ³-Voice at 75.4%
- Highest selectivity on the board at 66.5% — the best judgement about when speaking is warranted. [3]
- Interrupts the user 46.9% of the time, more than double every model above it bar one. High capability, talks over you. [3]
- No public API pricing. Access is gated behind a subscription, so cost per minute cannot be established at all. [3]
Sources
Every figure traces to one of these. Last verified 2026-09-18.
- OpenAI API pricing — all OpenAI token, per-minute and per-character rates.
- Gemini Developer API pricing — Gemini Live and TTS rates, and the dual per-minute figures used for the reconciliation check.
- τ³-Voice leaderboard (Sierra) — Overall pass^1 and all four interaction metrics. "Overall" excludes the Banking domain.
- xAI — Grok Voice Think Fast 2.0 — and reported $0.08/min pricing.
- Sierra — τ-voice benchmark methodology
- Artificial Analysis — speech-to-speech — independent quality index, conversational dynamics, Big Bench Audio, and the Qwen hourly rate.
- OpenAI — GPT-Live-1 in the API
- OpenAI developer forum — gpt-realtime-1.5 expressiveness regression
- Google — introducing Gemini 3.8 Live — the five benchmarks Google chose to cite.
- Deepgram — VAQI evaluation of gpt-realtime
- Artificial Analysis — Speech Agent Arena — blind-preference Elo, the fifth benchmark Google cited.
Prices change without notice and several of these models are in preview. Derived per-minute figures are arithmetic on published rates under the stated assumptions, not quoted prices — verify against the vendor before committing spend.