Photo Credit: Modulate

AI-powered startup Modulate, the company behind “audio-native voice intelligence model” Velma, has scored a $25 million raise.

Somerville-headquartered Modulate recently announced the multimillion-dollar round, which Steve Jurvetson-founded Future Ventures led. Meanwhile, existing backers Lakestar (which took the lead on Modulate’s $30 million Series A in August 2022) and Hyperplane participated as well.

Established in 2017, Modulate intends to use the newly obtained capital to “increase investment across AI/ML research” and build “new industry models” – while also growing its team and expanding its current products.

Chief among these products is the initially mentioned Velma, billed as “[t]he Ensemble Listening Model (ELM) architecture that’s 2x-4x more accurate than LLMs.”

As for this described accuracy’s benefits in practice, the moderation-focused system is said to monitor customer support lines, video game chats (“Rainbow Six Siege has reduced critical toxic voice chat by 50%”), and more in real time for various emotions, deepfakes, “disruptive behavior,” and a whole lot else.

Additionally, Modulate deals in transcription and offers a distinct “Velma Triage” product, according to its website.

Put differently, a significant portion of the business’s operations fall outside the industry at present. However, the relevant products are advertised as detecting “AI-generated music and synthetic singing” to boot, and as demonstrated by Universal Music-partnered ElevenLabs, a music-side expansion isn’t out of the question for well-funded AI voice players.

Furthermore, Modulate only debuted its AI music detection model in late June, indicating that the tool provides platforms with “segment-based probabilities for whether AI-generated vocals or…instrumentals are present continuously throughout a clip.” Interestingly, the product “is designed to generalize beyond any single AI music generator and to identify broader AI generation patterns,” per the announcement.

Back to the present, Modulate co-founder and CEO Carter Huffman framed voice as “a primary interface for AI” and touted his company’s “cost and compute” advantages over “traditional large models.”

“We’re already using audio-native AI to protect organizations from deepfake attacks, help voice agents understand emotion and respond with more empathy, identify dangerous behavior in online conversations, and monitor whether voice agents are actually performing the way they’re supposed to,” Huffman stated in part.

“Underneath all of that are more than a hundred specialized models working together to understand what’s really happening across audio, with dramatically less cost and compute than traditional large models,” he concluded.

Last month, the major labels participated in Stability AI’s $76 million Series B, and AI-powered “financial command center” Quantizr secured $5 million in seed capital.

On the generation side, Warner Music-, Believe-, and BMG-partnered Suno is still grappling with high-stakes infringement claims from (among others) Universal Music and Sony Music. At the same time, new AI products from Udio, Klay Vision, and Spotify alike are also forthcoming.