Uploaded September 2025 | Updated September 2026, 2 hours ago
Alibaba just dropped Qwen3-Omni, a 30B model that natively unifies text, image, audio, and video in one system—no trade-offs, no regressions, and even real-time streaming.
It supports 119 text languages, 19 speech inputs, and 10 outputs.
Plus, they open-sourced three flavors: Instruct (general tasks), Thinking (reasoning), and Captioner (low-hallucination audio descriptions). That last one is a big deal: it directly fills a gap in the community for reliable accessibility tools, showing how fine-tuning can fight hallucinations in practice.
With state-of-the-art performance on audio/AV benchmarks and built-in tool calling, this model is a serious push toward seamless multimodal AI.
Let me know which news I should cover next, and I’ll tag you!
I’m Louis-François, CTO & co-founder at Towards AI. Follow for tomorrow’s no-BS roundup 🚀
#AInews #Qwen3 #AlibabaAI #short
Alibaba just dropped Qwen3-Omni, a 30B model that natively unifies text, image, audio, and video in one system—no trade-offs, no regressions, and even real-time streaming.
It supports 119 text languages, 19 speech inputs, and 10 outputs.
Plus, they open-sourced three flavors: Instruct (general tasks), Thinking (reasoning), and Captioner (low-hallucination audio descriptions). That last one is a big deal: it directly fills a gap in the community for reliable accessibility tools, showing how fine-tuning can fight hallucinations in practice.
With state-of-the-art performance on audio/AV benchmarks and built-in tool calling, this model is a serious push toward seamless multimodal AI.
Let me know which news I should cover next, and I’ll tag you!
I’m Louis-François, CTO & co-founder at Towards AI. Follow for tomorrow’s no-BS roundup 🚀
#AInews #Qwen3 #AlibabaAI #short










