Uploaded August 2026 | Updated September 2026, 2 weeks ago
At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths.
Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task.
Apply to Y Combinator: ycombinator.com/apply
Work at a startup: ycombinator.com/jobs
At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Jitesh Jain about building video agents that can adapt their reasoning to videos of different lengths.
Current agents often struggle with long, open-ended video questions because temporal grounding is unreliable and training data is expensive. SAGE combines visual tools with transcripts and web search, then uses synthetic question-answer data, tool trajectories, and reinforcement learning to teach the model when each source of information is useful. As videos become longer, the agent takes more reasoning steps and improves more over the base model, suggesting that it is learning to spend effort according to the task.
Apply to Y Combinator: ycombinator.com/apply
Work at a startup: ycombinator.com/jobs










