Uploaded September 2023 | Updated September 2026, 2 days ago
Abstract: As models continue to rapidly evolve in complexity and scale, the status quo of how they are being evaluated and the quality of benchmarks has not significantly changed. This inertia leaves challenges in evaluation and data quality unaddressed, which results in the potential for erroneous conclusions. In this talk, I highlight these challenges in answering information-seeking questions. First, I discuss the failures of standard evaluation techniques such as lexical matching and other automated evaluation alternatives. Our study reveals the need for accurate evaluation mechanisms as we shift from extractive to generative models. Second, in the absence of culturally representative data in multilingual information retrieval, I introduce our effort to collect a new high-quality dataset to address this gap. Our dataset allows for a more realistic understanding of non-English retrieval settings. Next, I discuss our new dataset, collected via human-LLM collaboration, that paves the way for developing open-source generative search systems with attributable answers. Finally, with high-quality datasets at our disposal, I conclude by emphasizing the critical role of factuality and accessibility in shaping future helpful QA models.
Bio: Ehsan Kamalloo is a Post-doctoral Fellow working with Jimmy Lin at the David R. Cheriton School of Computer Science, University of Waterloo. He completed his PhD at the University of Alberta, focusing on robust knowledge acquisition in question answering. He is broadly interested in the intersection of natural language processing and information retrieval with a primary objective of trustworthy and reliable information access over massive unstructured text. His work has been published in top conferences and journals including ACL, CIKM, NAACL, SIGIR, TACL, and WWW.
Abstract: As models continue to rapidly evolve in complexity and scale, the status quo of how they are being evaluated and the quality of benchmarks has not significantly changed. This inertia leaves challenges in evaluation and data quality unaddressed, which results in the potential for erroneous conclusions. In this talk, I highlight these challenges in answering information-seeking questions. First, I discuss the failures of standard evaluation techniques such as lexical matching and other automated evaluation alternatives. Our study reveals the need for accurate evaluation mechanisms as we shift from extractive to generative models. Second, in the absence of culturally representative data in multilingual information retrieval, I introduce our effort to collect a new high-quality dataset to address this gap. Our dataset allows for a more realistic understanding of non-English retrieval settings. Next, I discuss our new dataset, collected via human-LLM collaboration, that paves the way for developing open-source generative search systems with attributable answers. Finally, with high-quality datasets at our disposal, I conclude by emphasizing the critical role of factuality and accessibility in shaping future helpful QA models.
Bio: Ehsan Kamalloo is a Post-doctoral Fellow working with Jimmy Lin at the David R. Cheriton School of Computer Science, University of Waterloo. He completed his PhD at the University of Alberta, focusing on robust knowledge acquisition in question answering. He is broadly interested in the intersection of natural language processing and information retrieval with a primary objective of trustworthy and reliable information access over massive unstructured text. His work has been published in top conferences and journals including ACL, CIKM, NAACL, SIGIR, TACL, and WWW.










