
September 2023 Releases
Within the first two seconds of any outbound call, an important classification has to be made: is this a person, a voicemail greeting, an IVR menu, or a call screener? That decision determines whether an AI agent gets deployed on the call at all. Routing an agent to a voicemail wastes money, but failing to identify a human, or even a call screener, costs you the chance to connect with a customer.
That is the operating reality of outbound voice AI within any business. Even with answering machine detection (AMD) running on every call, in high volume outbound campaigns with e.g., 5-10% answer rates, the industry average is that roughly half of the interactions routed to agents are still machines, not people.
That’s because most AI voice agent platforms and CCaaS rely on legacy AMD solutions like Twilio's, or a first-generation fine-tuned model like Bland's, to make that determination. The issue with these solutions is they either let too many answering machines through to AI agents, wasting minutes and reducing the ROI of outbound calling. Or they indiscriminately block call screeners, such as iOS's Call Screening feature, robbing your AI agents of live conversations they would have otherwise won.
With Regal’s next-generation Voicemail Detection model, you can now classify voicemail, iOS or Google call screeners as separate categories with very high accuracy rather than collapsing them into "machines." This categorization is what enables screener-specific handling, ultimately reducing missed agent connections.
.png)
Most AMD products (whether API-based or embedded in your CCaaS) publicly report a single blended “accuracy” number that doesn’t tell the whole story or let customers understand how the model will perform for their specific campaigns. But, not all call campaigns are the same. For example, progressive dialing campaigns have high call volumes of low intent leads, so naturally most responses are machines, not humans.
In that kind of environment, an AMD model can score well on overall accuracy by labeling most calls as “machines”, since the majority of these calls do not connect with humans. This overlabeling can still result in a strong overall accuracy score, even though it is quietly misclassifying humans or call screeners as machines, and dropping calls before a human ever gets a chance to connect.
Regal's Voicemail Detection model is multi-class. It classifies every call into one of five categories:
1. Automated Voicemail
2. Personal Recording
3. Google Screener
4. iOS Call Screener
5. Human Speech
Since Regal's model separates screeners and voicemail as distinct classes rather than collapsing them into "machine," it opens the door to things a binary classifier can't do: screener-specific handling and campaign-level funnel analytics. Additionally it allows us to report accuracy across each of these classes, and continue to train the model -- as inevitably new call screening behaviors, device-native features, and voicemail formats emerge across carriers and operating systems.
To train our Voicemail Detection model, we started with real outbound call audio from over 163,000 B2C calls and built a labeled dataset using a combination of AI-assisted labeling and manual human review. An LLM assessed transcripts and assigned categories at scale, while our team stepped in to hand-label the roughly 3,000 edge cases that automated labeling handles unreliably such as dead air and short ambiguous fragments.
Before committing to our five-class taxonomy, we ran a clustering analysis on audio embeddings to verify that each class was acoustically distinct, not just conceptually different. This gave us confidence that the model could actually learn to separate them, and highlighted where the boundaries were harder (live speech and personal voicemail greetings, for instance).

From there, we built on top of WavLM-Base-Plus*, an open-source audio representation model, training custom aggregation and classification layers on its embeddings using 2-second call clips. We tuned the model to prioritize human recall, since missing a real person costs you a conversation, while letting a voicemail through only costs agent minutes.
To evaluate performance, before any training began, we sealed 15% of the labeled data as a held-out test set (constructed to be representative across B2C industries and label distribution), so our accuracy numbers reflect how the model performs on audio it has never seen.

For a closer look at the methodology behind the model and how we’re continuing to evolve it, keep an eye out for an upcoming technical post from our engineering team.
Regal's Voicemail Detection model sends ~45% fewer voicemails to agents compared to incumbent models, directly reducing wasted agent minutes and cost. Of calls we route to an agent, only ~19% turn out to be machines; for Twilio AMD, that number is ~40%.
The gap is even starker on screener classification. Regal detects iOS and Google call screeners with 96% accuracy, while Twilio, at 43%, effectively treats them as machines, hanging up on conversations that could have been won. And we do all of this without sacrificing human recall: both models classify live human speech at ~96% accuracy, so the gains come with no tradeoff on the calls that matter most.
.gif)
Regal's model also makes its determination in 2 seconds, a full second faster than Twilio's minimum 3-second classification.
Taken together, these improvements directly translate to significantly better ROI on outbound AI voice agent campaigns: more live conversations reached, fewer minutes wasted.

_______________________________________________________________________________
*WavLM-Base-Plus by Microsoft (Chen et al., arXiv:2110.13900, 2021), licensed under CC BY-SA 3.0. License: https://github.com/microsoft/UniSpeech/blob/main/LICENSE. Model: https://huggingface.co/microsoft/wavlm-base-plus.
Attribution
REGAL IS NOT A LAW FIRM AND DOES NOT PROVIDE LEGAL SERVICES. DISTRIBUTION OF THIS LICENSE DOES NOT CREATE AN ATTORNEY-CLIENT RELATIONSHIP. REGAL PROVIDES THIS INFORMATION ON AN "AS-IS" BASIS. REGAL MAKES NO WARRANTIES REGARDING THE INFORMATION PROVIDED, AND DISCLAIMS LIABILITY FOR DAMAGES RESULTING FROM ITS USE
Ready to see Regal in action?
Book a personalized demo.



