VLAA Thinker: Dive Into Vision Language Models with Inherent Reasoning @arizeai
VLAA Thinker: Dive Into Vision Language Models with Inherent Reasoning  @arizeai
Uploaded July 2025 | Updated September 2026, 2 weeks ago
Reasoning models are quite popular these days, but most are limited to language. VLAA-Thinker offers vision-language reasoning data and training recipes for building a strong vision-language model with inherent reasoning ability. The building team open-sourced everything -- including the training data, code, and powerful model weights. VLAA-Thinker currently ranks at the top of the OpenCompass visual reasoning leaderboard.

In this session, one of the creators of VLAA-Thinker -- Haoqin Tu, PhD Student, UC Santa Cruz --- share details on dataset curation, insights from model training, and performance results.

⏱️ Chapters
0:00 - Introduction: Vision Language Models & Multimodal AI
1:00 - VLAA Thinker Overview: Training Data, Recipes, and Reasoning
2:30 - Prompting Evolution: From CoT to Inherent Reasoning
4:10 - Inheriting Reasoning: DeepSeek-R1 and Training Implications
5:30 - Research Questions: Is Two-Stage Training Necessary?
6:30 - SFT Pipeline and Data Generation Using DeepSeek
8:00 - How Rewriting and Verifiers Clean Reasoning Traces
9:00 - Pseudo Reasoning Problems from SFT
10:30 - RL-Only Training and Reward Design (GRPO)
12:00 - Mixed Reward Module: Math, IOU, and Open Reasoning
13:30 - SFT vs RL: Which Works Best?
15:00 - Leaderboard Results on Vision Reasoning Benchmarks
16:00 - Conclusion: Toward Transparent, Generalizable Multimodal Reasoning

More about VLAA Thinking family: github.com/UCSC-VLAA/VLAA-Thinking
Follow Haoqin Tu: https://x.com/haoqint
VLAA Thinker: Dive Into Vision Language Models with Inherent ReasoningPlan and Act: Enabling Agents to Solve Long Horizon TasksOne AI Question - why should you work at Arize, with Meredith MendeA Deep Dive Into Automated RAG Evaluation with open-rag-evalBenchmarking LLM Costs: GPT-5.5, Kimi K3, DeepSeek, and 8 More Models | AI BuildersA Watermark for Large Language ModelsGoogle TUMIX AI Agent Paper, Explained By Its AuthorOpenClaw vs Hermes: The Future of Open-Source AI Agents | Arize Observe 2026One AI Question - what do you do at night,  doom prompting with Matt WilsonAI Enablement At Enterprise Scale1.4 Billion Smiles: How PepsiCo Scales AI with PurposeProving a Prompt Fix Works in Production with Phoenixs PXI
Arize AI |

VLAA Thinker: Dive Into Vision Language Models with Inherent Reasoning

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER