Uploaded September 2025 | Updated September 2026, 2 hours ago
Forget bigger LLMs — Apple thinks the future of AI is faster, smaller, and in your pocket 👇
They just released FastVLM, a family of tiny vision-language models (0.5B, 1.5B, 7B) (process images + text and generate text).
Some highlights:
• 5x faster TTFT than SmolVLM, one of the fastest VLMs today
• Hybrid encoder (convolution + transformer) → efficient high-res processing with fewer, higher-quality tokens
• Great benchmark results, ideal for real-time apps
• Open-weights on Hugging Face (but research only, no commercial use)
• Limitation: input resolution fixed to square sizes → you’ll need to resize/pad images
TL;DR: Super interesting for researchers, less so for commercial apps — but a big step for small, on-device VLMs.
I’m Louis-François — PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#ai #apple #fastvlm #short
Forget bigger LLMs — Apple thinks the future of AI is faster, smaller, and in your pocket 👇
They just released FastVLM, a family of tiny vision-language models (0.5B, 1.5B, 7B) (process images + text and generate text).
Some highlights:
• 5x faster TTFT than SmolVLM, one of the fastest VLMs today
• Hybrid encoder (convolution + transformer) → efficient high-res processing with fewer, higher-quality tokens
• Great benchmark results, ideal for real-time apps
• Open-weights on Hugging Face (but research only, no commercial use)
• Limitation: input resolution fixed to square sizes → you’ll need to resize/pad images
TL;DR: Super interesting for researchers, less so for commercial apps — but a big step for small, on-device VLMs.
I’m Louis-François — PhD dropout, now CTO & co-founder at Towards AI. Follow me for tomorrow’s no-BS AI roundup 🚀
#ai #apple #fastvlm #short










