Fish Audio secures $52M seed round following $21M revenue milestone
Fish Audio has raised $52 million in a seed funding round led by Coreline Ventures and Capital Today to scale its generative AI voice models for creators and enterprise customers. The company currently generates $21 million in annual recurring revenue and provides synthetic voice APIs to partners including HeyGen and LiveKit.
Key Takeaways
- Seed funding co-led by Coreline Ventures and Capital Today included participation from 359 Capital and Parable.
- Fish Audio serves high-growth B2B partners including HeyGen for AI avatars and LiveKit for real-time voice agents.
- The startup's latest S2.1 Pro model is restricted to a paid API, departing from its early purely open-source roots.
- Automated takedown systems now verify and remove unauthorized voice clones in under three minutes to address creator concerns.
Why It Matters
Fish Audio's rapid ascent to $21 million in revenue signals a shift where high-fidelity, steerable voice AI is moving from a creative gimmick to critical enterprise infrastructure. By providing fine-grained controls via 15,000 natural language parameters, the company targets the latency and expressiveness gaps that have hindered voice agents in customer service and gaming. Its success with partners like HeyGen demonstrates a maturing supply chain for multimodal AI, where specialized vendors provide the underlying audio stack. Watch for the release of their audio-understanding model later this year, which aim to compete directly with integrated voice offerings from major labs like OpenAI and ElevenLabs.
Additional Context
The funding arrives as the generative audio sector sees massive capitalization and intensified competition for enterprise contracts. ElevenLabs recently reached a $1.1 billion valuation following a $500 million Series D led by Sequoia Capital, per Substack reporting in February 2026. ElevenLabs has since pivoted toward comprehensive conversational infrastructure, integrating with IBM watsonx Orchestrate in March 2026 to support 70 languages for government and financial service agents. This enterprise push is mirrored by infrastructure providers like LiveKit, which reported in early 2026 that over 200,000 developers are using its platform to build real-time voice applications for brands like Tesla.
Simultaneously, the regulatory environment is tightening for synthetic voice providers. According to the FCC in February 2026, AI-generated voices are strictly classified as 'artificial' under the Telephone Consumer Protection Act, making unsolicited AI-voice robocalls illegal and subject to fines of up to $1,500 per call. This legal landscape, combined with state-level protections like Tennessee’s ELVIS Act, has forced startups like Fish Audio to prioritize automated ownership verification. Per TechCrunch, the sector is also facing pressure from major labs; OpenAI updated its Advanced Voice Mode in June 2025 to include realistic cadence and emotional nuance, potentially threatening startups that rely solely on architectural patterns for enterprise AI agents as their primary differentiator.
Read full article at techcrunch.com
Enjoy our coverage?
Add StreamingMeme as a preferred source on Google to see more of our streaming news at the top of your Search results.
Add as preferred source