Linus ✦ Ekenstam
@LinusEkenstam
You're not ready for this 🚨
AI voices can now laugh, whisper and sigh on command
Fish Audio just dropped S2.1 Pro, the most expressive voice AI I've heard. More emotional range than ElevenLabs and Cartesia
6x more affordable than ElevenLabs 💸
Here's the full breakdown 👇
For years, voice AI had one goal: sound human. That race is over. Everything sounds realistic now.
The new benchmark is expressiveness. Emotion, nuance, timing. Can a voice *perform* a line, not just read it?
That's what S2.1 Pro was built for.
I ran the same emotional script through Fish Audio, ElevenLabs and Cartesia. Same words, completely different performances. The pauses, the intonation, the way Fish Audio's voice breathes between lines. It's not close. Watch the video below.
The wild part is how you control it. You direct it like a voice actor, in plain text:
[nervous laugh] → expressive laughing.
[cry] → it cries
[long sigh] → it expels a long sigh
Plus word-level control over pronunciation, emphasis and pacing. Generating speech is out. Directing performances is in.
And it's not just for voiceovers. S2.1 Pro streams at ~90ms to first audio. That's fast enough for live conversation, with rhythm that survives interruptions and topic changes. It's what makes voice agents finally feel human.
The practical stuff:
• Clone a voice from 15 seconds of audio
• 80+ languages in one model
• Open-source roots, open weights models, self-hostable
• 1/6 of the cost of ElevenLabs
HeyGen already integrated Fish Audio. More will follow.
Try it: https://t.co/DugUDKVjfm (there's a free tier)
Reply with a line you'd love to hear an AI voice truly perform, and I'll run the best ones through S2.1 Pro and post the results 👇
347 likes 304.7K views
THIS DEVELOPER JUST KILLED THE VOICE CLONING INDUSTRY WITH ONE GITHUB REPO.
Claude can now speak in your own voice across 23 languages through a single MCP call.
Free. Local. Offline.
What used to cost $22 to $330/month with ElevenLabs is now open source under the MIT license.
Jamie Pine just dropped voicebox on GitHub, and it already has over 41k stars.
It clones any voice from just 10 seconds of audio, lets you choose from seven different TTS engines, and runs entirely on your own machine, so your voice and recordings never leave your computer.
Connect it to Claude Code, Cursor, Cline, or any MCP compatible agent with a single voicebox.speak call, and your AI can start replying in your cloned voice.
It also adds global voice dictation that works in any app with one hotkey and can generate speech in 23 languages, including Arabic, Japanese, and Polish.
One open source app now replaces both ElevenLabs and Wispr Flow.
If you’ve been waiting to give your AI agent a voice, this is it.
Repo below 👇
3.8K likes 270.9K views