Sourced pricing · verified 2026-09-18

What a minute of voice actually costs

Every realtime speech-to-speech model on Sierra's τ³-Voice leaderboard, plus the newer releases it hasn't rated yet — priced on both axes, with the arithmetic shown and the gaps left as gaps.

4

different billing units in one market: per audio token, per minute, per hour, per character. No vendor publishes all of them.

$0.05/min

gpt-live-1 ranks #1 at 81.7% — and is billed flat per minute. It has no token price at all.

80×

OpenAI's cached audio input ($0.40/1M) against uncached ($32.00/1M). Any single $/min figure is meaningless without stating this.

2

models sit on the efficient frontier — and Gemini 3.8 Live lands on the cheaper one's exact price, with no τ³-Voice score published to place it.

The finding that shapes this page: You want price per minute and price per 1M tokens. Most vendors publish exactly one of the two. Where a vendor publishes a token price and a documented token-per-second rate, the per-minute figure here is derived and marked as such. Where a vendor bills flat per minute and never discloses a token rate — now true of the top three models on the leaderboard — the per-1M-token cell is not applicable, not estimated. Filling those cells would mean inventing numbers.

Assumptions you can change

A “minute of conversation” is not one thing. These two controls drive every derived figure.

Reconciliation check — passing. Google is the only vendor publishing both axes, which makes it a test of the formula rather than just an input. Gemini Live bills audio at 25 tokens/sec (not the 32 tok/s rate that applies to uploaded audio files). 1500 tok/min × $12.00/1M = $0.0180/min — matching Google's own published $0.018/min exactly. OpenAI's rates (1 token/100ms in, 1 token/50ms out) give $0.019/min in and $0.077/min out for a full talking minute on gpt-realtime-2.1, matching independently reported figures.

Flat per-minute models don't move when you drag this. Token-billed ones do — output audio is the expensive half.

Input audio caching

Applies only where a vendor publishes a cached audio rate. Google publishes none for Live audio, so Gemini rows are unaffected.

The formula, applied identically to every token-billed row
$/min = (user_seconds × in_tok_per_sec × $in / 1e6) + (model_seconds × out_tok_per_sec × $out / 1e6)
0%20%40%60%80%$0.01$0.02$0.05$0.10COST PER MINUTE → (LOG)← τ³-VOICE PASS^1gemini-3.8-live — same price as the frontierreleased Sep 15 · not yet rated heregpt-live-1grok-think-fast-1.0grok-think-fast-2.0 · highgemini-3.1-flash-live · highgpt-realtime-2 · xhighgpt-realtime-2 · minimalgpt-realtime-1.5gpt-realtime-1.0gemini-3.1-flash-live · minimalgemini-2.5-native-audioNO PUBLIC PRICEPine Voice Previewqwen3.5-omni-realtimegrok-voice-fast-1.0Cascaded baseline · ASR + LLM + TTS

Violet marks sit on the efficient frontier — nothing is both cheaper and better. Hollow marks in the right-hand rail score on the benchmark but publish no usable price, so they cannot be placed on a price axis at all.

Priced, but not yet rated on τ³-Voice Overall

These have a real, computable cost per minute and no score on this benchmark's Overall table. They are the mirror image of the hollow marks in the chart's right-hand rail, which have a score and no price.

gemini-3.8-live-extended-thinking$0.0112/min82.6 AA quality (#1) · 68.6% τ-Voice agentic · 35.1% Sierra Banking · 97.7% Big Bench Audio
gemini-3.8-live$0.0112/min2nd on Speech Agent Arena · 97 languages, switched mid-conversation
gemini-2.5-flash-native-audio$0.0112/min
gpt-realtime-2.1-mini$0.0150/min
gpt-audio-mini$0.0150/min
gpt-realtime-2.1$0.0480/min
gpt-audio$0.0480/min

Realtime speech-to-speech

Ranked by τ³-Voice Overall pass^1. Interaction metrics are the leaderboard's own measurements of conversational behaviour — the closest thing to a reproducible “does it feel human” number.

#ModelBilling$/1M audio in$/1M audio out$/minPass^1LatencyInterruptsRespons.
#1gpt-live-1
backend gpt-6 astra (medium)
OpenAI · Sep 10, 2026
per minuten/an/a$0.0500
published
81.7%2.54s22.0%95.5%
#2Pine Voice Preview
reasoning enabled
Pine AI · Aug 10, 2026
not publishedn/an/an/a75.4%2.14s46.9%92.7%
#3grok-voice-think-fast-1.0
reasoning enabled
xAI · Apr 23, 2026
per minuten/an/a$0.0500
published
67.3%1.21s19.7%100.0%
#4grok-voice-think-fast-2.0
high
xAI · Jul 29, 2026
per minuten/an/a$0.0800
published
62.5%1.71s10.7%100.0%
#5qwen3.5-omni-plus-realtime
Qwen · Mar 30, 2026
not publishedn/an/an/a53.7%1.70s36.1%99.2%
#6gemini-3.1-flash-live-preview
thinking high
Google · Mar 26, 2026
per token$3.00$12.00$0.0112
derived
43.8%3.15s18.9%84.7%
#7gpt-realtime-2
xhigh
OpenAI · May 7, 2026
per token$32.00
cached $0.40
$64.00$0.0480
derived
42.4%1.98s20.8%95.1%
#8gpt-realtime-2
minimal
OpenAI · May 7, 2026
per token$32.00
cached $0.40
$64.00$0.0480
derived
38.5%1.44s18.7%99.8%
#9grok-voice-fast-1.0
xAI · Dec 17, 2025
not publishedn/an/an/a38.3%1.15s84.3%91.3%
#10gpt-realtime-1.5
OpenAI · Feb 23, 2026
per token$32.00
cached $0.40
$64.00$0.0480
derived
35.3%1.39s13.5%99.9%
#11Cascaded baseline
ASR + LLM + TTS
Multiple · May 19, 2026
not publishedn/an/an/a31.2%4.24s64.4%79.1%
#12gpt-realtime-1.0
OpenAI · Aug 28, 2025
per token$32.00
cached $0.40
$64.00$0.0480
derived
30.4%1.53s12.3%97.8%
#13gemini-3.1-flash-live-preview
thinking minimal
Google · Mar 26, 2026
per token$3.00$12.00$0.0112
derived
28.6%1.64s9.8%97.0%
#14gemini-live-2.5-flash-native-audio
Google · Dec 12, 2025
per token$3.00$12.00$0.0112
derived
25.8%1.43s20.6%80.6%
gemini-3.8-live-extended-thinkingnewest
multi-step reasoning; speaks while thinking
Google · Sep 15, 2026
Vendor-published, other benchmarks: 82.6 AA quality (#1) · 68.6% τ-Voice agentic · 35.1% Sierra Banking · 97.7% Big Bench Audio
per token$3.00$12.00$0.0112
derived
not ratedn/an/an/a
gemini-3.8-livenewest
built for scale and cost efficiency
Google · Sep 15, 2026
Vendor-published, other benchmarks: 2nd on Speech Agent Arena · 97 languages, switched mid-conversation
per token$3.00$12.00$0.0112
derived
not ratedn/an/an/a
gpt-realtime-2.1
not yet on τ³-Voice
OpenAI · 2026
per token$32.00
cached $0.40
$64.00$0.0480
derived
not ratedn/an/an/a
gpt-realtime-2.1-mini
not yet on τ³-Voice
OpenAI · 2026
per token$10.00
cached $0.30
$20.00$0.0150
derived
not ratedn/an/an/a
gpt-audio
not realtime; not on τ³-Voice
OpenAI · 2026
per token$32.00$64.00$0.0480
derived
not ratedn/an/an/a
gpt-audio-mini
not realtime; not on τ³-Voice
OpenAI · 2026
per token$10.00$20.00$0.0150
derived
not ratedn/an/an/a
gemini-2.5-flash-native-audio
not yet on τ³-Voice
Google · Dec 2025
per token$3.00$12.00$0.0112
derived
not ratedn/an/an/a

Reading the interaction columns: Latency is time to respond (lower better). Interrupts is how often the model talks over the user (lower better). Responsiveness is whether it responds at all when it should (higher better). n/a means the vendor publishes nothing on that axis.

Everyone else, on Google's benchmarks

Google cited five benchmarks for its newest model and none of them was τ³-Voice Overall. Here is every model scored on each one Google chose — plus, last, the one it left out.

Speech-to-Speech Quality Index
Google cited
Artificial Analysis · composite of 4 components, 0–100
Gemini 3.8 Live Ext Thinking (High)
82.6
GPT-Live-1 (Astra, medium)
81.5
Grok Voice Think Fast 2.0 High
81.3
Gemini 3.8 Live
76.0
Qwen Audio 3.0 Realtime Plus
66.8
Google #1, by 1.1 points over GPT-Live-1 (Astra, medium). A model needs results for all four components to receive an index score at all, which is why some high performers carry none.
τ-Voice agentic task completion
Google cited
Artificial Analysis · share of tasks completed
Gemini 3.8 Live Ext Thinking
68.6%
GPT-Live-1 Astra
67.9%
Grok Voice Think Fast 2.0
56.5%
Qwen Audio 3.0 Realtime Plus
54.6%
Google #1, by 0.7 points over GPT-Live-1 Astra. A margin this narrow sits inside almost any error bar.
Big Bench Audio (speech reasoning)
Google cited
Artificial Analysis · reasoning questions delivered as audio
Qwen Audio 3.0 Realtime Plus
99.2%
Gemini 3.8 Live Ext Thinking
97.7%
Grok Voice Think Fast 2.0
97.2%
GPT-Live-1
90.1%
Google places 2nd of 4, behind Qwen Audio 3.0 Realtime Plus (99.2%). Measures reasoning delivered as audio, not conversational skill — a model can lead one and trail the other.
τ-Voice · Banking domain
Google cited
Sierra · long policy tasks; explicitly excluded from Overall
Gemini 3.8 Live Ext Thinking
35.1%
GPT-Live-1 Astra
32.0%
xAI-Realtime
16.5%
Google #1, by 3.1 points over GPT-Live-1 Astra. Long policy tasks, and the domain Sierra explicitly excludes from Overall.
Conversational Dynamics
Sub-metric
Artificial Analysis · turn-taking, pauses, interruptions
StepAudio 3 Realtime
98.9%
Qwen Audio 3.0 Realtime Plus
98.4%
Grok Voice Think Fast 2.0
95.1%
GPT-Live-1 Astra
94.9%
Gemini 3.8 Live Ext Thinking
91.9%
A component of an index Google cited, never quoted on its own. Google places 5th of 5, behind StepAudio 3 Realtime (98.9%). Turn-taking, pauses and interruptions — the metric closest to whether it feels like a conversation.
τ³-Voice Overall
Not cited
Sierra · the leaderboard the rest of this page is built on
gpt-live-1
81.7%
Pine Voice Preview
75.4%
grok-voice-think-fast-1.0
67.3%
grok-voice-think-fast-2.0
62.5%
qwen3.5-omni-plus-realtime
53.7%
gemini-3.1-flash-live (high)
43.8%
Not cited by Google. Google places 6th of 6, behind gpt-live-1 (81.7%). The leaderboard the rest of this page is built on. No Gemini 3.8 entry exists here; Google is represented by the 3.1 predecessor.

Bars start at zero, which is why most of them look alike — that is the finding, not a rendering fault. Purple bars are Google's. The last panel is the benchmark Google did not quote.

Model names differ by source and are printed exactly as each source publishes them. Artificial Analysis rates Gemini 3.8 Live Extended Thinking; Sierra's Overall table has no 3.8 entry at all, so Google's best-placed model there is the 3.1 Flash Live predecessor. Those are different models and are not stacked into one bar. Separately, GPT-Live-1 appears as 67.9% on Artificial Analysis and 81.7% on Sierra — same model, different harness. That 13.8-point gap is why nothing here is merged onto a shared axis.

One-way models: speech out, speech in

Kept in a separate table on purpose. These bill per character or per minute of audio and do not hold a conversation, so putting them on the same axis as the models above would be a false comparison.

ModelProviderJobBilled inRateText in / 1M
tts-1OpenAIText to speech1M characters$15.00
tts-1-hdOpenAIText to speech1M characters$30.00
gpt-4o-mini-ttsOpenAIText to speech1M audio out tokens$12.00$0.60
gemini-2.5-flash-preview-ttsGoogleText to speech1M audio out tokens$10.00$0.50
gemini-3.1-flash-tts-previewGoogleText to speech1M audio out tokens$20.00$1.00
gemini-2.5-pro-preview-ttsGoogleText to speech1M audio out tokens$20.00$1.00
gpt-transcribeOpenAISpeech to textminute of audio$0.0045
gpt-4o-mini-transcribeOpenAISpeech to textminute of audio$0.0030$1.25
whisperOpenAISpeech to textminute of audio$0.0060
gpt-4o-transcribeOpenAISpeech to textminute of audio$0.0060$2.50
gpt-4o-transcribe-diarizeOpenAISpeech to text + speakersminute of audio$0.0060$2.50
gpt-live-transcribeOpenAIRealtime transcriptionminute of audio$0.0170
gpt-realtime-whisperOpenAIRealtime transcriptionminute of audio$0.0170
gpt-realtime-translateOpenAIRealtime translationminute of audio$0.0340

How they actually sound

What people report about quality and lifelikeness. That is opinion, and it is kept structurally separate from every number above.

Opinion — not measured

Reported reactions from reviews, vendor posts and developer forums. Every claim carries a numbered source. The measured numbers sit under each entry for contrast.

gpt-live-1

OpenAI · #1 on τ³-Voice at 81.7%

  • Reviewers describe it as dropping the "robotic turn-taking" of earlier voice AI — it listens and speaks at once rather than trading turns. [7]
  • Language-learning app Speak reported wrong interruptions cut by roughly 80% after switching. [7]
  • Scores 97.3% on Artificial Analysis' conversational-dynamics measure, the highest reported. [6]
  • A launch-day tally put reaction at 78% positive across 482 responses — critics named latency and occasional wrong answers. [7]
  • Its measured 2.54s latency is the second-slowest of any model on the leaderboard, despite the naturalness praise. [3]
Measured: Latency 2.54s · Interrupts 22.0% · Responsiveness 95.5%

Gemini 3.8 Live / Live Extended Thinking

Google · released Sep 15, 2026 · absent from τ³-Voice; predecessor scores 43.8%

  • Uses early verbal cues — acknowledging with something like "let me check that" — specifically so it doesn't read as a system waiting to answer. [9]
  • Extended Thinking reasons and speaks simultaneously, which is the architectural claim developers responded to rather than a demo trick. [9]
  • Google reports #1 on Artificial Analysis' Speech-to-Speech Quality Index at 82.6, and 97.7% on Big Bench Audio. [9]
  • Detects and switches between 97 languages mid-conversation — a lifelikeness axis nothing else here claims. [9]
  • Every score Google published is on a benchmark other than τ³-Voice Overall, so it cannot be ranked against the models above it. [9]
  • Developers flag latency gaps and patchy regional availability, and want speech-specific benchmarks before trusting headline numbers. [9]
  • Its rated predecessor is the slowest model on the board at 3.15s with the second-lowest responsiveness — unproven that 3.8 fixes this. [3]
Measured: No τ³-Voice measurement exists. Predecessor (3.1 Flash Live, high): Latency 3.15s · Interrupts 18.9% · Responsiveness 84.7%

grok-voice-think-fast-2.0

xAI · 62.5% on τ³-Voice

  • Tuned toward real conversational habits: shorter sentences, one question at a time, less filler. [4]
  • The only top-five model to break one second on Artificial Analysis, cutting latency from 1.25s to 0.70s. [6]
  • Users report graceful handling of background noise mid-sentence rather than derailing. [4]
  • Priced 60% above Think Fast 1.0 ($0.05 → $0.08/min) while scoring lower than 1.0 on this benchmark. [3][4]
  • Selectivity drops to 35.9% from 1.0's 51.5% — it is less discriminating about when to speak. [3]
Measured: Latency 1.71s · Interrupts 10.7% · Responsiveness 100.0%

gpt-realtime family

OpenAI · 25.8–42.4% on τ³-Voice

  • Instruction-followable delivery — tone and pacing can be directed in the prompt. [1]
  • Deepgram's independent VAQI evaluation gives it the best miss-rate in its test set. [10]
  • A developer-forum thread titles the 1.5 release a "major regression in voice expressiveness", reporting accents largely gone. [8]
  • Practitioners building outbound voice agents describe a prosody plateau — acceptable for low-stakes assistants, not for callers who did not opt in. [10]
  • One engineer's writeup of building on the Realtime API is titled around the voices sounding robotic, and moves to a hybrid stack. [10]
Measured: gpt-realtime-2 (xhigh): Latency 1.98s · Interrupts 20.8% · Responsiveness 95.1%

Pine Voice Preview

Pine AI · #2 on τ³-Voice at 75.4%

  • Highest selectivity on the board at 66.5% — the best judgement about when speaking is warranted. [3]
  • Interrupts the user 46.9% of the time, more than double every model above it bar one. High capability, talks over you. [3]
  • No public API pricing. Access is gated behind a subscription, so cost per minute cannot be established at all. [3]
Measured: Latency 2.14s · Interrupts 46.9% · Selectivity 66.5%

Sources

Every figure traces to one of these. Last verified 2026-09-18.

  1. OpenAI API pricing — all OpenAI token, per-minute and per-character rates.
  2. Gemini Developer API pricing — Gemini Live and TTS rates, and the dual per-minute figures used for the reconciliation check.
  3. τ³-Voice leaderboard (Sierra) — Overall pass^1 and all four interaction metrics. "Overall" excludes the Banking domain.
  4. xAI — Grok Voice Think Fast 2.0 — and reported $0.08/min pricing.
  5. Sierra — τ-voice benchmark methodology
  6. Artificial Analysis — speech-to-speech — independent quality index, conversational dynamics, Big Bench Audio, and the Qwen hourly rate.
  7. OpenAI — GPT-Live-1 in the API
  8. OpenAI developer forum — gpt-realtime-1.5 expressiveness regression
  9. Google — introducing Gemini 3.8 Live — the five benchmarks Google chose to cite.
  10. Deepgram — VAQI evaluation of gpt-realtime
  11. Artificial Analysis — Speech Agent Arena — blind-preference Elo, the fifth benchmark Google cited.

Prices change without notice and several of these models are in preview. Derived per-minute figures are arithmetic on published rates under the stated assumptions, not quoted prices — verify against the vendor before committing spend.