Uploaded August 2026 | Updated September 2026, 2 weeks ago
At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Alexander Braverman about a test-time scaling method for improving reasoning in smaller language models.
Instead of relying on a larger teacher model or an external reward model, the method uses the language model’s own self-critique as a heuristic in an A*-inspired search. It explores multiple reasoning paths, deprioritizes weaker branches, and searches for a stronger answer using the same underlying model. On mathematical reasoning benchmarks, the method improved accuracy more efficiently than other test-time approaches at comparable token and runtime budgets.
Apply to Y Combinator: ycombinator.com/apply
Work at a startup: ycombinator.com/jobs
At our inaugural YCML at Startup School, YC Partner Ankit Gupta speaks with Alexander Braverman about a test-time scaling method for improving reasoning in smaller language models.
Instead of relying on a larger teacher model or an external reward model, the method uses the language model’s own self-critique as a heuristic in an A*-inspired search. It explores multiple reasoning paths, deprioritizes weaker branches, and searches for a stronger answer using the same underlying model. On mathematical reasoning benchmarks, the method improved accuracy more efficiently than other test-time approaches at comparable token and runtime budgets.
Apply to Y Combinator: ycombinator.com/apply
Work at a startup: ycombinator.com/jobs










