Uploaded April 2025 | Updated September 2026, 1 day ago
Speaker: Kexin Huang, PhD Student, Stanford University
Abstract: (provisional)
Hypotheses are central to information acquisition, decision-making, and discovery. However,many real-world hypotheses are abstract, highlevel statements that are difficult to validate directly. This challenge is further intensified bythe rise of hypothesis generation from Large Language Models (LLMs), which are prone to hallucination and produce hypotheses in volumes thatmake manual validation impractical. Here wepropose POPPER, an agentic framework for rigorous automated validation of free-form hypotheses.Guided by Karl Popper’s principle of falsification, POPPER validates a hypothesis using LLMagents that design and execute falsification experiments targeting its measurable implications. Anovel sequential testing framework ensures strictType-I error control while actively gathering evidence from diverse observations, whether drawnfrom existing data or newly conducted procedures.We demonstrate POPPER on six domains including biology, economics, and sociology. POPPERdelivers robust error control, high power, and scalability. Furthermore, compared to human scientists, POPPER achieved comparable performancein validating complex biological hypotheses whilereducing time by 10 folds, providing a scalable,rigorous solution for hypothesis validation.
Bio:
Kexin Huang (kexinhuang.com/) is a fourth-year PhD student in Computer Science at Stanford University, advised by Prof. Jure Leskovec. His research focuses on leveraging AI to drive novel, deployable, and interpretable biomedical discoveries, while also tackling fundamental AI challenges such as multi-modal modeling, uncertainty quantification, and agentic reasoning. His work has been published in Nature Medicine, Nature Biotechnology, Nature Chemical Biology, Nature Biomedical Engineering, and machine learning conferences including NeurIPS, ICML, ICLR, and UAI. His research has been featured in major media outlets such as Forbes, WIRED, and MIT Technology Review. He has also contributed to machine learning research at leading companies and institutions, including Genentech, GSK, Pfizer, IQVIA, Flatiron Health, Dana-Farber Cancer Institute, and Rockefeller University.
Speaker: Kexin Huang, PhD Student, Stanford University
Abstract: (provisional)
Hypotheses are central to information acquisition, decision-making, and discovery. However,many real-world hypotheses are abstract, highlevel statements that are difficult to validate directly. This challenge is further intensified bythe rise of hypothesis generation from Large Language Models (LLMs), which are prone to hallucination and produce hypotheses in volumes thatmake manual validation impractical. Here wepropose POPPER, an agentic framework for rigorous automated validation of free-form hypotheses.Guided by Karl Popper’s principle of falsification, POPPER validates a hypothesis using LLMagents that design and execute falsification experiments targeting its measurable implications. Anovel sequential testing framework ensures strictType-I error control while actively gathering evidence from diverse observations, whether drawnfrom existing data or newly conducted procedures.We demonstrate POPPER on six domains including biology, economics, and sociology. POPPERdelivers robust error control, high power, and scalability. Furthermore, compared to human scientists, POPPER achieved comparable performancein validating complex biological hypotheses whilereducing time by 10 folds, providing a scalable,rigorous solution for hypothesis validation.
Bio:
Kexin Huang (kexinhuang.com/) is a fourth-year PhD student in Computer Science at Stanford University, advised by Prof. Jure Leskovec. His research focuses on leveraging AI to drive novel, deployable, and interpretable biomedical discoveries, while also tackling fundamental AI challenges such as multi-modal modeling, uncertainty quantification, and agentic reasoning. His work has been published in Nature Medicine, Nature Biotechnology, Nature Chemical Biology, Nature Biomedical Engineering, and machine learning conferences including NeurIPS, ICML, ICLR, and UAI. His research has been featured in major media outlets such as Forbes, WIRED, and MIT Technology Review. He has also contributed to machine learning research at leading companies and institutions, including Genentech, GSK, Pfizer, IQVIA, Flatiron Health, Dana-Farber Cancer Institute, and Rockefeller University.

![From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2
AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy!, but the rich variety of standardized exams has remained a landmark challenge. Even as recently as 2016, the best AI system could achieve merely 59.3% on an 8th Grade science exam.
This talk reports success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exams non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern Natural Language Processing (NLP) methods can result in mastery on this task. While not a full solution to general question-answering (the questions are limited to 8th Grade multiple-choice science) it represents a significant milestone for the field. [ Paper at AI Magazine 41 (4), Winter 2020, https://arxiv.org/pdf/1909.01958.pdf ] From F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2](https://i.ytimg.com/vi/CR3aICkhCJM/mqdefault.jpg)








