Uploaded July 2025 | Updated September 2026, 2 weeks ago
Reasoning models are quite popular these days, but most are limited to language. VLAA-Thinker offers vision-language reasoning data and training recipes for building a strong vision-language model with inherent reasoning ability. The building team open-sourced everything -- including the training data, code, and powerful model weights. VLAA-Thinker currently ranks at the top of the OpenCompass visual reasoning leaderboard.
In this session, one of the creators of VLAA-Thinker -- Haoqin Tu, PhD Student, UC Santa Cruz --- share details on dataset curation, insights from model training, and performance results.
⏱️ Chapters
0:00 - Introduction: Vision Language Models & Multimodal AI
1:00 - VLAA Thinker Overview: Training Data, Recipes, and Reasoning
2:30 - Prompting Evolution: From CoT to Inherent Reasoning
4:10 - Inheriting Reasoning: DeepSeek-R1 and Training Implications
5:30 - Research Questions: Is Two-Stage Training Necessary?
6:30 - SFT Pipeline and Data Generation Using DeepSeek
8:00 - How Rewriting and Verifiers Clean Reasoning Traces
9:00 - Pseudo Reasoning Problems from SFT
10:30 - RL-Only Training and Reward Design (GRPO)
12:00 - Mixed Reward Module: Math, IOU, and Open Reasoning
13:30 - SFT vs RL: Which Works Best?
15:00 - Leaderboard Results on Vision Reasoning Benchmarks
16:00 - Conclusion: Toward Transparent, Generalizable Multimodal Reasoning
More about VLAA Thinking family: github.com/UCSC-VLAA/VLAA-Thinking
Follow Haoqin Tu: https://x.com/haoqint
Reasoning models are quite popular these days, but most are limited to language. VLAA-Thinker offers vision-language reasoning data and training recipes for building a strong vision-language model with inherent reasoning ability. The building team open-sourced everything -- including the training data, code, and powerful model weights. VLAA-Thinker currently ranks at the top of the OpenCompass visual reasoning leaderboard.
In this session, one of the creators of VLAA-Thinker -- Haoqin Tu, PhD Student, UC Santa Cruz --- share details on dataset curation, insights from model training, and performance results.
⏱️ Chapters
0:00 - Introduction: Vision Language Models & Multimodal AI
1:00 - VLAA Thinker Overview: Training Data, Recipes, and Reasoning
2:30 - Prompting Evolution: From CoT to Inherent Reasoning
4:10 - Inheriting Reasoning: DeepSeek-R1 and Training Implications
5:30 - Research Questions: Is Two-Stage Training Necessary?
6:30 - SFT Pipeline and Data Generation Using DeepSeek
8:00 - How Rewriting and Verifiers Clean Reasoning Traces
9:00 - Pseudo Reasoning Problems from SFT
10:30 - RL-Only Training and Reward Design (GRPO)
12:00 - Mixed Reward Module: Math, IOU, and Open Reasoning
13:30 - SFT vs RL: Which Works Best?
15:00 - Leaderboard Results on Vision Reasoning Benchmarks
16:00 - Conclusion: Toward Transparent, Generalizable Multimodal Reasoning
More about VLAA Thinking family: github.com/UCSC-VLAA/VLAA-Thinking
Follow Haoqin Tu: https://x.com/haoqint










