Evaluating and Enhancing Language Model Factuality @allenai
Evaluating and Enhancing Language Model Factuality  @allenai
Uploaded March 2025 | Updated September 2026, 1 day ago
Abstract: Language models (LMs) are increasingly adopted in real-world applications, yet their tendency to generate factual errors remains a major concern. In this talk, I will describe my work on LM factuality, i.e., its consistency with established facts. I address factuality challenges across two key dimensions: evaluation and enhancement. On the evaluation front, I will present a factuality evaluation framework comprising an updatable benchmark curated from real-world LM usage and a fine-grained evaluation technique that robustly identifies LM inaccuracies. For factuality enhancement, I will propose two complementary approaches: (1) a post-processing framework that verifies and refines LM outputs against external knowledge sources; and (2) learnable intervention systems that leverages LMs' internal representations of truth to adjust generations at inference time. Together, these methods advance our understanding of factuality challenges and offer practical pathways to improve LM reliability.

Bio: Farima Fatahi Bayat is a Ph.D. candidate in the Computer Science and Engineering Department at University of Michigan, advised by Prof. H. Jagadish and Prof. Lu Wang. Her research focuses on advancing responsible AI, with a particular emphasis on enhancing the factuality of Language Models (LMs). Her recent works include creating evaluation benchmarks to assess LMs’ factuality, designing adaptive intervention frameworks that enable uncertainty expression, and building correction mechanisms to increase the quality of LM output.
Evaluating and Enhancing Language Model Factuality🤝 Molmo meets SAMIntegrated Systems for Computational Scientific DiscoverySoft Robotics and AI towards Embodied Intelligence | Embodied AI Lecture Series at AI2Data Leverage: A Framework for Empowering the Public and Mitigating Harms of AI | AI2Robot Learning by Understanding Egocentric VideosDR Tulu | An open, end-to-end training recipe for long-form deep researchStructure Modeling in Language Models❓ Molmo AMA: Your Questions Answered!Simulation and Generalization in VLA Models for Robotic Manipulation🥽 Molmo Vision Pro Demo - Augmenting how we see with AIReliable Evaluation and High-Quality Data: Building Blocks for Helpful Question Answering Systems
Ai2 |

Evaluating and Enhancing Language Model Factuality

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER