Modulate Raises $25M—Voice AI Must Prove It Hears More Than Words

|Author: QUASA Editorial Team|5 min read| 2
Modulate Raises $25M—Voice AI Must Prove It Hears More Than Words

On September 28, 2026, Boston-based Modulate announced $25 million in new funding led by Future Ventures, with Hyperplane and Lakestar participating, bringing its disclosed funding total to $60 million; CEO Carter Huffman called voice “a primary interface for AI.” The money is intended to expand the company’s audio models, developer tools, partnerships and hiring.

The investment backs Velma, Modulate’s platform for analyzing the sound of conversations as well as their words. Its proposed uses include spotting synthetic voices and suspected fraud, assessing voice agents and detecting harmful behavior in online voice chat. The central commercial question is whether those signals remain useful on real calls, where a missed warning and a false alarm can carry different costs.

The round backs an audio layer for enterprise voice

Velma’s premise is that a transcript preserves spoken words but loses acoustic information. Tone, emphasis and characteristics of synthetic speech remain in the recording; they cannot be reconstructed from text alone. Modulate’s Ensemble Listening Model architecture combines specialized models that extract signals from audio and use them to identify events relevant to a particular application.

That architecture gives the company a route into several kinds of voice system. A fraud application might need an alert during a suspected impersonation call. A voice-agent operator might need to know whether an automated conversation left a customer dissatisfied or broke a rule. Trust and safety teams face a different task when assessing harmful behavior in group voice chat. These uses share an audio input, but the decision and the cost of an error differ in each setting.

The expansion plan includes new software development kits, APIs and models tailored to industries, along with partner integrations and additional deployment options. Hiring is planned across research, engineering and developer relations. These are plans for broader access and capability, rather than evidence that the forthcoming tools are already available to developers.

Audio volume and deepfake accuracy are company disclosures

SiliconANGLE’s account gives the company’s figures as more than 10 million hours of audio processed each month, more than 600 million hours processed in total, more than 100 specialized models, and 98.9% deepfake-detection accuracy on public benchmark data. The figures describe different things: workload, system design and performance on a test. Independent coverage of the financing does not independently measure that workload or reproduce the accuracy result.

The audio-hour totals indicate substantial use of Modulate’s systems, but they do not disclose how many paying enterprises generated the traffic, how concentrated it is among customers or what revenue it produced. The model count explains how the platform is assembled; it does not measure how well the full system handles a particular customer’s calls. Neither figure can substitute for a task-specific outcome.

The deepfake percentage has a narrower meaning than reliability in a live fraud workflow. Overall accuracy combines correct decisions on genuine and synthetic samples in a particular test. For incoming calls, a fraud team would also need to know how often the detector misses an imitation and how often it flags a genuine speaker. Those rates can shift with the sample mix, background noise, new cloning methods and the threshold chosen for an alert.

A public transcription result answers a different question

The Open ASR results file lists Modulate’s multilingual entry at a 3.84125 average word-error rate across its English short-form tests, while Zoom’s scribe_v2_pro entry is lower at 3.5925 in the same file. Lower word-error rate means fewer differences between a transcription and its reference text. This is a publicly visible Modulate result, though the listed comparison does not put its entry first.

That result matters because recognizing words is part of understanding a conversation. It measures transcription on defined datasets, however, rather than synthetic-voice detection, caller intent or voice-agent performance. A leaderboard position can also change as entries are added. The visible transcription row and the company’s deepfake accuracy figure therefore cannot be treated as interchangeable proof of Velma’s broader capabilities.

Timing introduces another distinction. An accurate transcript delivered after a call may help with review or coaching; an impersonation warning must arrive early enough to affect a live interaction. Delay, missed detections and false alarms all shape whether an audio signal is useful in production, even when a model performs well on a public test.

The next evidence will come from specific voice tasks

For fraud teams, an alert could trigger additional authentication, while a false alarm could interrupt a legitimate caller. For voice-agent operators, a calm exchange can still end with an unresolved request, so a simple positive-or-negative reading of tone may miss the outcome that matters. Online safety teams likewise need detections that distinguish harmful conduct from ordinary speech in the context of a conversation.

Those differences explain the investor appeal of a reusable audio layer and the difficulty of proving one. The funding gives Modulate resources to put more tools in developers’ hands. As its planned SDKs and industry models reach them, task-specific results on realistic calls—including missed detections, false alarms and alert timing—will show how far its listening capabilities extend beyond transcription and disclosed usage.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0