Reliable Evaluation and High-Quality Data: Building Blocks for Helpful Question Answering Systems @allenai
Reliable Evaluation and High-Quality Data: Building Blocks for Helpful Question Answering Systems  @allenai
Uploaded September 2023 | Updated September 2026, 2 days ago
Abstract: As models continue to rapidly evolve in complexity and scale, the status quo of how they are being evaluated and the quality of benchmarks has not significantly changed. This inertia leaves challenges in evaluation and data quality unaddressed, which results in the potential for erroneous conclusions. In this talk, I highlight these challenges in answering information-seeking questions. First, I discuss the failures of standard evaluation techniques such as lexical matching and other automated evaluation alternatives. Our study reveals the need for accurate evaluation mechanisms as we shift from extractive to generative models. Second, in the absence of culturally representative data in multilingual information retrieval, I introduce our effort to collect a new high-quality dataset to address this gap. Our dataset allows for a more realistic understanding of non-English retrieval settings. Next, I discuss our new dataset, collected via human-LLM collaboration, that paves the way for developing open-source generative search systems with attributable answers. Finally, with high-quality datasets at our disposal, I conclude by emphasizing the critical role of factuality and accessibility in shaping future helpful QA models.


Bio: Ehsan Kamalloo is a Post-doctoral Fellow working with Jimmy Lin at the David R. Cheriton School of Computer Science, University of Waterloo. He completed his PhD at the University of Alberta, focusing on robust knowledge acquisition in question answering. He is broadly interested in the intersection of natural language processing and information retrieval with a primary objective of trustworthy and reliable information access over massive unstructured text. His work has been published in top conferences and journals including ACL, CIKM, NAACL, SIGIR, TACL, and WWW.
Reliable Evaluation and High-Quality Data: Building Blocks for Helpful Question Answering SystemsOlmo 3 | A family of leading fully open LMs and complete model flowDiscrete diffusion with planned denoisingMolmo 2 | Robotics ApplicationsThe Future is Hear: Advances in Computational Audition and Sound ManipulationSkill it! A Data-Driven Skills Framework for Understanding and Training Language ModelsAsta | an agentic ecosystem that advances scientific discoveryOpenBot: Turning Smartphones into Robots | Embodied AI Lecture Series at AI2Rethinking LLM efficiencyWere bringing training text out in the open, Introducing OLMoTraceSide Effects May Include Homogenization and Overusing ClichesGooAQ: Open Question Answering with Diverse Answer Types | AI2
Ai2 |

Reliable Evaluation and High-Quality Data: Building Blocks for Helpful Question Answering Systems

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER