top of page

Fish Audio Raised $52M in Seed Funding

  • Writer: Karan Bhatia
    Karan Bhatia
  • 1 day ago
  • 2 min read

Fish Audio, the most expressive, emotionally controllable real-time voice model, led by Rissa Cao, Shijia Liao, and Jiahua Liu, has announced $52 million in seed funding on its first anniversary, led by Coreline Ventures and Capital Today with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphaist Partners, and angels.


Traditional text-to-speech systems have focused on generating accurate words but often fail to capture natural delivery, emotion, and expression. Drawing from experience building voice AI at Amazon Alexa and Meta, Fish Audio is taking a different approach, building a voice layer designed not just to speak text, but to deliver more expressive and human-like audio experiences across applications.


Fish Audio’s foundation was shaped by research into making AI voices more expressive and natural. Chief Scientist Shijia Liao began exploring new voice architectures while at NVIDIA, experimenting with models trained on a single GPU. His early work on a dual autoregressive architecture became the foundation for S2, a technology designed to advance AI voice generation beyond traditional text-to-speech systems.


In just one year, Fish Audio has rapidly scaled its platform and technology. The company has launched five state-of-the-art audio models, grown its team from 3 to 22 employees, reached $21M in annual recurring revenue, and attracted more than 8 million creators, developers, and enterprises. Its community now includes over 2 million voice models, while S2.1 Pro has achieved strong preference in blind listening tests, supporting 83+ languages with advanced emotion and cadence control.


Fish Audio’s models have been shaped by real-world user interactions, using community feedback to improve performance across languages, accents, and edge cases. The company credits creators, developers, and communities such as game developers and anime fans for helping transform Fish Speech from an open-source project into a widely used voice AI platform. With a focus on accessibility and openness, Fish Audio has built a global community of more than 8 million users.


Fish Audio is seeing growing adoption among enterprises, which now represent two-thirds of its revenue alongside developers. The company provides enterprise-grade voice AI with features including on-premise deployment, zero-data-retention policies, and HIPAA-compliant configurations. Its S2.1 Pro model combines high-quality voice generation with fast inference, multilingual performance, and natural delivery, helping businesses solve the challenge of creating AI voices that sound truly expressive.


Fish Audio sees text-to-speech as the first step toward a broader audio intelligence platform. The company is expanding beyond TTS into Audio Language Models (Audio LM) and speech-to-speech technologies, while investing in developer tools, APIs, enterprise capabilities, and partnerships to bring more expressive AI voice experiences to a wider range of applications.


To mark the milestone, Fish Audio is making its S2.1 Pro model available free to developers through its API until the end of August, while offering discounts on creator plans and migration incentives for users switching from other providers. The company attributes this strategy to improvements in its inference infrastructure, including custom FP8 kernels that enable 8,000+ tokens per second on a single H200, significantly reducing serving costs.


𝐑𝐞𝐦𝐨𝐭𝐞 𝐇𝐢𝐫𝐞 𝐖𝐢𝐭𝐡 𝐔𝐬: 𝐌𝐞𝐧𝐥𝐨 𝐓𝐚𝐥𝐞𝐧𝐭 helps technology companies hire exceptional remote talent from India. Build your team with carefully curated engineers, product leaders, designers, GTM professionals, and more roles. Learn More At: https://www.menlotimes.com/menlo-talent

Menlo Times is a global media platform covering AI, Deeptech, Venture Capital, Fintech, Robotics, and Security through news, analysis, and insights from founders and operators.
  • Instagram
  • Facebook
  • X(Formerly Twitter)
  • LinkedIn
  • YouTube
© 2026 Menlo Times. All rights reserved.
bottom of page