Molmo 2 | Counting objects and actions @allenai
Molmo 2 | Counting objects and actions  @allenai
Uploaded December 2025 | Updated September 2026, 4 hours ago
Researcher Zixin Ma demonstrates Molmo 2's counting capabilities.

Molmo 2 is a state-of-the-art open multimodal model suite capable of precise spatial and temporal understanding of video, image, and multi-image sets. Building on the global impact of Molmo, which pioneered image pointing for multimodal AI systems, Molmo 2 introduces breakthrough capabilities in video pointing, multi-frame reasoning, and object tracking.

Learn more: allenai.org/molmo
Discuss in Discord: discord.gg/ai2
Molmo 2 | Counting objects and actionsBLADE: Benchmarking Language Model Agents for Data-Driven ScienceJust-DREAM-about-it: Figurative Language Understanding with DREAM-FLUTEHelping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward HackingTowards the Age of Computation and AI for High Performance ClimateWhen Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric MemoriesOpen-Ended Learning Leads to Generally Capable Agents | Embodied AI Lecture Series at AI2From LLMs to Agents: Generalizability from the Inside OutData-Centric Approaches to Adapting Foundation ModelsThe University of Washington eScience Institute: a Home for Data-Intensive DiscoveryGeneralization for Robot Learning In The Wild | Embodied AI Lecture series at AI2Towards Generalist Agents for Accelerating Scientific Discovery171
Ai2 |

Molmo 2 | Counting objects and actions

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER