Automated Hypothesis Validation with Agentic Sequential Falsifications @allenai
Automated Hypothesis Validation with Agentic Sequential Falsifications  @allenai
Uploaded April 2025 | Updated September 2026, 1 day ago
Speaker: Kexin Huang, PhD Student, Stanford University

Abstract: (provisional)
Hypotheses are central to information acquisition, decision-making, and discovery. However,many real-world hypotheses are abstract, highlevel statements that are difficult to validate directly. This challenge is further intensified bythe rise of hypothesis generation from Large Language Models (LLMs), which are prone to hallucination and produce hypotheses in volumes thatmake manual validation impractical. Here wepropose POPPER, an agentic framework for rigorous automated validation of free-form hypotheses.Guided by Karl Popper’s principle of falsification, POPPER validates a hypothesis using LLMagents that design and execute falsification experiments targeting its measurable implications. Anovel sequential testing framework ensures strictType-I error control while actively gathering evidence from diverse observations, whether drawnfrom existing data or newly conducted procedures.We demonstrate POPPER on six domains including biology, economics, and sociology. POPPERdelivers robust error control, high power, and scalability. Furthermore, compared to human scientists, POPPER achieved comparable performancein validating complex biological hypotheses whilereducing time by 10 folds, providing a scalable,rigorous solution for hypothesis validation.

Bio:
Kexin Huang (kexinhuang.com/) is a fourth-year PhD student in Computer Science at Stanford University, advised by Prof. Jure Leskovec. His research focuses on leveraging AI to drive novel, deployable, and interpretable biomedical discoveries, while also tackling fundamental AI challenges such as multi-modal modeling, uncertainty quantification, and agentic reasoning. His work has been published in Nature Medicine, Nature Biotechnology, Nature Chemical Biology, Nature Biomedical Engineering, and machine learning conferences including NeurIPS, ICML, ICLR, and UAI. His research has been featured in major media outlets such as Forbes, WIRED, and MIT Technology Review. He has also contributed to machine learning research at leading companies and institutions, including Genentech, GSK, Pfizer, IQVIA, Flatiron Health, Dana-Farber Cancer Institute, and Rockefeller University.
Automated Hypothesis Validation with Agentic Sequential FalsificationsOpenWebMath: An Open Dataset of High-Quality Mathematical Web TextFrom F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2Meet an AI Using ScholarReliability and interactive debugging for language modelsUsing Asta AutoDiscovery: AI-powered autonomous scientific discoveryRobots Need To Reduce, Reuse, and Recycle | Embodied AI Lecture series at AI2Cross-Task Generalization via Natural Language Crowdsourcing InstructionsAI Scaffolding Systems for the Academic Peer Review EcosystemInference-Time Policy Customization Through Interactive Task SpecificationMolmoWeb Inference LibraryPrompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts
Ai2 |

Automated Hypothesis Validation with Agentic Sequential Falsifications

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER