Fish Audio, a startup based in Palo Alto, has raised $52 million to expand its library of natural language controls for AI-generated voices. The company offers both open-source and paid models catering to creators and enterprises alike, aiming to provide more expressive and steerable options than currently available on the market.
Their latest model, S2.1 Pro, is only available through their paid API. Fish Audio’s CEO Rissa Cao highlights that different businesses have unique needs: realism for AI avatars, expressiveness in gaming characters, and natural-sounding voices for customer service and sales operations.
One of the challenges Fish Audio faced was ensuring user consent when training models on submitted voices. The company has since automated a takedown process to address concerns more efficiently. However, this doesn’t stop unauthorized uploads from persisting until reported by the rightful owner.
The crowded market for AI voice generation includes well-funded competitors like ElevenLabs and WellSaid. Fish Audio believes its fine-grained controls and cost-efficient model training will give them a competitive edge. They plan to release an audio understanding model this year and develop a speech-to-speech model in the future.







