Uploaded December 2025 | Updated September 2026, 1 day ago
Researcher Yue Yang demonstrates Molmo 2's ability to reason across multiple documents and images at once.
Molmo 2 is a state-of-the-art open multimodal model suite capable of precise spatial and temporal understanding of video, image, and multi-image sets. Building on the global impact of Molmo, which pioneered image pointing for multimodal AI systems, Molmo 2 introduces breakthrough capabilities in video pointing, multi-frame reasoning, and object tracking.
Learn more: allenai.org/molmo
Discuss in Discord: discord.gg/ai2
Researcher Yue Yang demonstrates Molmo 2's ability to reason across multiple documents and images at once.
Molmo 2 is a state-of-the-art open multimodal model suite capable of precise spatial and temporal understanding of video, image, and multi-image sets. Building on the global impact of Molmo, which pioneered image pointing for multimodal AI systems, Molmo 2 introduces breakthrough capabilities in video pointing, multi-frame reasoning, and object tracking.
Learn more: allenai.org/molmo
Discuss in Discord: discord.gg/ai2










