VYPR
researchPublished Sep 29, 2026· 1 source

Modulate Secures $25 Million to Combat Deepfake Voices with Advanced Audio AI

Modulate has raised $25 million to enhance its audio AI platform, designed to detect deepfake voices and secure automated voice communications against fraud and impersonation.

Modulate, a company specializing in audio-native AI, has successfully secured $25 million in a funding round led by Future Ventures, with participation from Hyperplane and Lakestar. This latest investment brings the company's total funding to $60 million and will be instrumental in accelerating its research and development in artificial intelligence, expanding its engineering team, enhancing developer tools, forging new partnerships, and broadening deployment options for organizations integrating voice intelligence into their security and communication infrastructures.

The infusion of capital comes at a critical juncture as the proliferation of convincing synthetic speech technologies renders traditional telephone-based security cues increasingly unreliable. Modulate's innovative approach moves beyond merely analyzing transcripts of conversations. Instead, its platform delves into the nuances of how something is spoken, examining factors such as emotion, tone, emphasis, intent, conversational patterns, and crucially, detecting signs of artificial generation. This granular analysis is vital for combating sophisticated threats like vishing (voice phishing) and executive impersonation attacks, where the spoken words might seem innocuous, but the delivery—marked by urgency, manipulation, or an unnatural voice—reveals the underlying danger.

At the core of Modulate's offering is Velma, a real-time conversation-understanding platform engineered to identify a range of problematic events. These include fraud attempts, instances of harassment, expressions of customer dissatisfaction, violations of company policy, and failures in automated voice agent interactions. The platform's efficacy is powered by its Ensemble Listening Model (ELM), which orchestrates over 100 specialized audio models. This architecture contrasts with traditional methods that rely on a single, monolithic foundation model, offering significant advantages in efficiency and resource utilization.

Modulate claims that its ELM design can achieve up to 1,000 times greater inference efficiency compared to single-model approaches. This translates to substantially reduced requirements for computing power, memory, and energy consumption during audio analysis. The deepfake detection capability, a key security feature, has demonstrated impressive performance. As of August 19, 2026, Modulate reported an average equal error rate of 1.104% across 14 Speech DF Arena datasets, indicating a detection accuracy of approximately 98.9% and securing a top ranking on that benchmark at the time.

The company's models are currently processing over 10 million hours of audio monthly and have cumulatively analyzed more than 600 million hours. Beyond deepfake detection, Modulate has also achieved recognition for its transcription accuracy, ranking first among 88 evaluated models on Hugging Face's Open ASR Leaderboard in July. The company offers tiered pricing for its services, with batch transcription starting at $0.03 per hour and deepfake detection priced at $0.25 per hour.

With the new funding, Modulate plans to expand its API and SDK offerings, develop specialized models tailored for various industries, integrate with more partner platforms, and enhance its support for diverse deployment environments. The company is targeting critical sectors such as fraud prevention, healthcare security, contact center oversight, moderation for gaming and social platforms, child safety initiatives, and the supervision of AI voice agents.

While Modulate's technology offers powerful real-time analysis capabilities that could enable defenders to challenge suspicious callers or escalate risky sessions, the company cautions that its benchmark performance may not perfectly translate to all real-world scenarios. Factors like compressed telephone audio, background noise, unfamiliar languages, replay attacks, and novel voice generation techniques can impact accuracy. Security teams are advised to treat the technology as a supplementary detection layer rather than a definitive identity verification tool, and to tune detection thresholds based on their specific operational context.

As voice increasingly becomes a primary interface for AI interactions, the challenges of authentication, manipulation detection, and securing automated agents become paramount. Modulate's strategic focus on audio understanding positions it to address these emerging security needs, though a layered security approach combining audio AI with multifactor verification, transaction controls, and human oversight remains essential to mitigate risks effectively.

Synthesized by Vypr AI
Modulate Secures $25 Million to Combat Deepfake Voices with Advanced Audio AI · VYPR