Molmo 2 | Dense Captioning @allenai
Molmo 2 | Dense Captioning  @allenai
Uploaded December 2025 | Updated September 2026, 2 days ago
Researcher Jieyu Zhang demonstrates Molmo 2's dense captioning capabilities.

Molmo 2 is a state-of-the-art open multimodal model suite capable of precise spatial and temporal understanding of video, image, and multi-image sets. Building on the global impact of Molmo, which pioneered image pointing for multimodal AI systems, Molmo 2 introduces breakthrough capabilities in video pointing, multi-frame reasoning, and object tracking.

Learn more: allenai.org/molmo
Discuss in Discord: discord.gg/ai2
Molmo 2 | Dense CaptioningStudying Large Language Model Generalization with Influence FunctionsHere is Tülu 3 405B 🐫Explaining Answers with Entailment TreesOlmo 3 | Livestream with Hugging FaceDeduplication of Large-scale Text Datasets for Pretraining of Language ModelsRobot Learning with Sparsity and ScarcityTowards Data-Driven Scientific Discovery with Generative AI: From Mathematical Modeling to LLMsWildDet3D - an open model for monocular 3D detectionDeepEarth: Multimodal Probabilistic World Model with 4D Spacetime EmbeddingEnhancing Reasoning in Smaller Models through Self-TrainingMaking Health Knowledge Accessible Through Personalized Language Processing
Ai2 |

Molmo 2 | Dense Captioning

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER