Uploaded October 2025 | Updated September 2026, 2 hours ago
Qwen3-VL just dropped—and it’s a big one. 🚀
4B, 8B, and even a 30B MoE multimodal model capable of seeing, reasoning, and thinking across text, images, and even video. This series fuses vision and language seamlessly, handles 1M-token context, and lets you toggle “thinking mode” for deeper reasoning without swapping models.
Dense for predictable edge apps, MoE for cloud-scale inference. FP8 quantization means near–BF16 performance with tiny memory footprints.
And yes—expanded OCR in 32 languages.
This could redefine local multimodal agents and research workflows alike.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#Qwen3VL #MultimodalAI #AIResearch #short
Qwen3-VL just dropped—and it’s a big one. 🚀
4B, 8B, and even a 30B MoE multimodal model capable of seeing, reasoning, and thinking across text, images, and even video. This series fuses vision and language seamlessly, handles 1M-token context, and lets you toggle “thinking mode” for deeper reasoning without swapping models.
Dense for predictable edge apps, MoE for cloud-scale inference. FP8 quantization means near–BF16 performance with tiny memory footprints.
And yes—expanded OCR in 32 languages.
This could redefine local multimodal agents and research workflows alike.
I’m Louis-François, PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#Qwen3VL #MultimodalAI #AIResearch #short










