Uploaded September 2025 | Updated September 2026, 15 minutes ago
Microsoft just released a major breakthrough for AI voice tech — and it’s open source! 🎙️
They dropped VibeVoice, an open-source TTS framework for long-form, multi-speaker audio.
• Two models:
– VibeVoice-1.5B: 64K context, ~90 min
– VibeVoice-Large: ~10B params, 32K context, ~45 min
• Handles up to 4 speakers → perfect for podcasts or multi-voice audiobooks
• MIT-licensed → free for commercial use
• Outperforms baselines in realism, richness, and speaker similarity
• Still only supports English + Chinese for now
• Doesn’t handle background music or effects — pure speech only
• A smaller streaming model is also coming soon!
It’s not replacing things like GPT’s real-time voice API, but it’s a seriously powerful alternative for creating high-quality long-form audio.
Would you use this for podcasts or audiobooks? 👀
I’m Louis-François — PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀 #short
Microsoft just released a major breakthrough for AI voice tech — and it’s open source! 🎙️
They dropped VibeVoice, an open-source TTS framework for long-form, multi-speaker audio.
• Two models:
– VibeVoice-1.5B: 64K context, ~90 min
– VibeVoice-Large: ~10B params, 32K context, ~45 min
• Handles up to 4 speakers → perfect for podcasts or multi-voice audiobooks
• MIT-licensed → free for commercial use
• Outperforms baselines in realism, richness, and speaker similarity
• Still only supports English + Chinese for now
• Doesn’t handle background music or effects — pure speech only
• A smaller streaming model is also coming soon!
It’s not replacing things like GPT’s real-time voice API, but it’s a seriously powerful alternative for creating high-quality long-form audio.
Would you use this for podcasts or audiobooks? 👀
I’m Louis-François — PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀 #short










