Uploaded October 2025 | Updated September 2026, 6 hours ago
Scientific discovery has historically required the creative synthesis of paradigm-questioning insights, empirical observation, and mathematical
rigor—from Copernicus's hypothesis to Kepler's models to Newton's universal laws. Today, we stand at the threshold of automating this discovery
process using artificial intelligence, particularly large language models. However, current approaches often excel at recombining existing
knowledge while struggling with the creative skepticism and mathematical precision needed for genuine breakthroughs. This talk explores how
data-driven mathematical modeling serves as both a rigorous testbed and foundational component for reliable AI systems for scientific discovery.
Kazem discusses recent developments in leveraging language models for scientific model discovery, multi-modal approaches that combine symbolic
and data-driven representation learning, iterative scientific model discovery via programming with large language models, and novel benchmark
methodologies specifically designed to distinguish real discovery from memorization of existing scientific knowledge. Looking forward, he
outlines potential future directions including enhancing LLM creativity and paradigm-questioning capabilities, developing multi-modal capabilities
for direct integration of empirical data, and enhancing evaluation frameworks through collaboration with domain experts to validate authentic
scientific contributions and translate AI-generated hypotheses into scientific publications and breakthroughs.
Kazem is an AI Researcher at Capital One AI Foundations and a recent PhD graduate from Carnegie Mellon University, advised by Amir Barati Farimani. His research interests revolve around language models and AI-driven scientific discovery. His work explores leveraging large language models for
symbolic reasoning and mathematical model generation, developing AI systems that can discover and verify scientific models from data. During his PhD, he held research internships at Netflix Research and Electronic Arts, working on foundation models and large language model applications.
Scientific discovery has historically required the creative synthesis of paradigm-questioning insights, empirical observation, and mathematical
rigor—from Copernicus's hypothesis to Kepler's models to Newton's universal laws. Today, we stand at the threshold of automating this discovery
process using artificial intelligence, particularly large language models. However, current approaches often excel at recombining existing
knowledge while struggling with the creative skepticism and mathematical precision needed for genuine breakthroughs. This talk explores how
data-driven mathematical modeling serves as both a rigorous testbed and foundational component for reliable AI systems for scientific discovery.
Kazem discusses recent developments in leveraging language models for scientific model discovery, multi-modal approaches that combine symbolic
and data-driven representation learning, iterative scientific model discovery via programming with large language models, and novel benchmark
methodologies specifically designed to distinguish real discovery from memorization of existing scientific knowledge. Looking forward, he
outlines potential future directions including enhancing LLM creativity and paradigm-questioning capabilities, developing multi-modal capabilities
for direct integration of empirical data, and enhancing evaluation frameworks through collaboration with domain experts to validate authentic
scientific contributions and translate AI-generated hypotheses into scientific publications and breakthroughs.
Kazem is an AI Researcher at Capital One AI Foundations and a recent PhD graduate from Carnegie Mellon University, advised by Amir Barati Farimani. His research interests revolve around language models and AI-driven scientific discovery. His work explores leveraging large language models for
symbolic reasoning and mathematical model generation, developing AI systems that can discover and verify scientific models from data. During his PhD, he held research internships at Netflix Research and Electronic Arts, working on foundation models and large language model applications.


![Enhancing Reasoning in Smaller Models through Self-Training
Abstract: Smaller language models can develop robust reasoning capabilities through pre-training, fine-tuning, or knowledge distillation from large language models (LLMs). However, unlike LLMs that employ a diverse array of reasoning strategies, smaller models typically rely on a single dominant approach. This limitation restricts their effectiveness in handling different multi-step reasoning tasks, which require a wide range of strategies in order to solve them. To address this challenge, self-training leverages the model’s own generated data, enabling smaller models to autonomously learn and adapt their reasoning strategies for improved performance across diverse tasks.
I will talk about a self-guided iterative distillation framework (SIKeD [1]), which combines multi-strategy outputs from LLMs with self-generated data from the smaller model to identify the most effective strategy for a given task in an on-policy manner.
Later, I will talk about how this self-training approach can be extended to improve refinement in models, where a model can learn to iteratively refine its output, eventually learning to pick the right strategy in its first attempt (SMART [2]).
[1] https://arxiv.org/abs/2410.18574
[2] https://arxiv.org/abs/2410.16128
Bio: Kumar Shridhar is a final-year Ph.D. candidate at ETH Zürich, Switzerland, under the supervision of Prof. Mrinmaya Sachan from ETH and Dr. Nicholas Monath from Google DeepMind. Prior to his doctoral studies, he spent summers interning at FAIR, Microsoft Research, and Alexa AI, and improving conversational AI at different startups.
His research focuses on advancing the reasoning capabilities of large language models (LLMs) and developing efficient distillation methods to impart these skills to smaller models. Moreover he is also working model alignment, autonomous agents, and model refinement. He is also a member of Swiss AI initiative, where the team is training foundational models across various domains. Enhancing Reasoning in Smaller Models through Self-Training](https://i.ytimg.com/vi/SS59gCT2KKs/mqdefault.jpg)







