| Model | Best for | Voice CX Score | Time to first token | Cost per 1K turns |
|---|
| GPT-6 Sol | Best overall and regulated conversations | 786 | 870 ms | $19.88 |
| GPT-4.1 | Task completion and customer support | 785 | 793 ms | $25.34 |
| GLM 5.3 Flash | Open-source, personalization, and tool calling | 767 | 984 ms | $1.34 |
| GPT-6 Luna | Lowest cost per turn | 735 | 754 ms | $0.99 |
| GLM 4.7 | Fastest response time | 732 | 584 ms | $4.44 |
| GPT-4.1 Mini | Outbound calls | 764 | 810 ms | $5.18 |
| Gemini 3.1 Flash Live | Realtime speech-to-speech | 752 | 1,005 ms | $15.96 |
How We Ranked These Models
Every ranking in this article comes from the Regal Voice CX Benchmark, which tests how LLMs perform inside live voice conversations. We ran 27 models from OpenAI, Google, Anthropic, and open-source providers through the same 80 inbound and outbound call scenarios, producing 33,615 conversation turns and 316,311 scored samples. Scenarios cover customer support and lead qualification, on both single-state and multi-state agents.
Every turn is scored on five dimensions:
- Task progression: does the turn move the call toward its goal?
- Latency: median time to first token, meaning how long the caller waits before the agent starts responding.
- Accuracy and safety: is the agent correct, compliant, careful with sensitive data, and reliable with tools?
- Clarity and coherence: is the response easy to follow when it's spoken aloud?
- Personalization and engagement: does the agent adapt to the caller and keep the call moving?
These combine into a single Voice CX Score out of 1,000, weighted most heavily toward task progression and latency, because those two best predict whether a model holds up in production.
How each turn is graded: before scoring, every turn is classified by what it's trying to do: greeting, identity verification, knowledge answer, objection handling, appointment scheduling, and 14 other turn types. The type decides which of 29 underlying metrics apply. For example, hallucination and knowledge-base faithfulness are checked on knowledge answers, and PII leakage and verification rigor on identity checks, so a greeting is never marked down for failing a knowledge check.
Why latency matters so much: the 10 slowest models (1.3 seconds or more to first token) actually have slightly higher task progression than the 16 fastest (0.570 vs. 0.555). Even so, they average about 128 points lower on the Voice CX Score (623 vs. 751). On a phone call, two seconds of silence sounds like a dropped line, which is why flagship models like Claude Opus 5 (#21) and Gemini 3.1 Pro (#22) rank near the bottom. For the engineering side of this, see how Regal cut AI agent latency by 26%.
1. GPT-6 Sol — Best Overall LLM for Voice AI
GPT-6 Sol has the highest Voice CX Score in the benchmark, 786 out of 1,000, with a median time to first token of 870 ms. It's the only flagship model fast enough to finish first. It also leads on accuracy and safety (0.802) and clarity and coherence (0.880), ranks first on inbound calls (809) and lead qualification (795), and has the strongest guardrail compliance of any model tested. That makes it the safest pick for regulated conversations in insurance, lending, and healthcare.
The tradeoff: at $19.88 per 1,000 conversation turns, it costs roughly 15x more than GLM 5.3 Flash or GPT-5.6 Luna, both of which score within 20 points of it.
2. GPT-4.1 — Best for Task Completion
GPT-4.1 ranks #2 overall (785) and has the highest task progression score of all 27 models (0.626). That means it's the best at moving a call toward its goal, whether that's booking the appointment, qualifying the lead, or collecting the payment. It's also one of the fastest leaders at 793 ms, and it ranks first on customer support calls (791).
The tradeoff: it's the most expensive model in the top five at $25.34 per 1,000 turns, and its accuracy and safety score (0.771) trails GPT-6 Sol's.
3. GLM 5.3 Flash — Best Open-Source LLM for Voice AI
GLM 5.3 Flash ranks #3 overall (767), ahead of every Google and Anthropic model tested, at about $1.34 per 1,000 turns. It leads the benchmark on personalization and engagement (0.867) and tool calling (0.816), and it's effectively tied with GPT-4.1 on customer support calls (790 vs. 791).
The tradeoff: its task progression (0.546) is noticeably lower than the leaders', and its 984 ms response time is near the upper limit callers tolerate. Self-hosting or running it through an inference provider adds operational overhead.
4. GPT-6 Luna — Best Value LLM for Voice AI
GPT-6 Luna is the cheapest model in the benchmark at $0.99 per 1,000 turns, yet it still ranks #13 overall (735) with a fast 754 ms response time. It delivers the most Voice CX Score per dollar of any model tested, which suits high-volume programs like reminders, payment nudges, and re-engagement.
The tradeoff: its task progression (0.500) is low, so it's better for short, simple calls. If you need top-five quality on a budget, GPT-5.6 Luna ranks #4 (766) for $1.18 per 1,000 turns.
5. GLM 4.7 — Fastest LLM for Voice AI
GLM 4.7 has the fastest median time to first token of any model tested, at 584 ms. Claude Haiku 4.5 (662 ms) and Gemini 2.5 Flash (718 ms) are next. It ranks #15 overall (732) at $4.44 per 1,000 turns, and it's the pick when response speed is your hard constraint.
The tradeoff: it gives up task progression (0.527) for speed, which keeps it out of the top 10. Among the overall leaders, GPT-4.1 (793 ms) and GPT-4.1 Mini (810 ms) are the fastest.
6. GPT-4.1 Mini — Best for Outbound Calls
GPT-4.1 Mini has the highest outbound Voice CX Score (754), ahead of GPT-6 Sol (749) and GLM 5.3 Flash (747). Outbound is harder for most models because the agent has to open the call, earn attention, and drive the conversation itself. GPT-4.1 Mini handles it with the highest outbound task progression among the leaders (0.546), an 810 ms response time, and a cost of $5.18 per 1,000 turns. It ranks #5 overall (764).
The tradeoff: it has the lowest clarity (0.825) and personalization (0.827) scores in the top five, so conversations can feel more scripted.
7. Gemini 3.1 Flash Live — Best Realtime (Speech-to-Speech) Model
Gemini 3.1 Flash Live ranks #8 overall (752). It's the only realtime model in the top 10, and it leads its category by nearly 100 points. If your architecture streams audio directly to the model instead of running a speech-to-text, LLM, and text-to-speech pipeline, it's the model to beat. It also ranks third on customer support calls (776).
The tradeoff: being speech-native doesn't automatically mean better results. The other realtime models tested rank #19, #23, and #26, and GPT Realtime 2.1 costs $93.24 per 1,000 turns.
Schedule a demo with Regal