Welcoming AI as a New Colleague: How Should We Evaluate AI for Science? @allenai
Welcoming AI as a New Colleague: How Should We Evaluate AI for Science?  @allenai
Uploaded March 2026 | Updated September 2026, 3 minutes ago
AI is reshaping the work of researchers in real time. With the emergence of tools such as AI Scientist in 2024, a new era of AI-driven scientific discovery has begun. At the same time, high-profile venues are beginning to experiment with AI-assisted peer review. Yet the core question remains unresolved: How should we meaningfully evaluate AI for scientific work?

In this talk, Iryna highlights challenges encountered in several projects focused on AI-assisted scientific communication and discuss how these challenges inform evaluation practices. She examines methods for assessing AI-generated related-work sections and explore how automated reviewers can be evaluated with respect to research-design reasoning and novelty assessment. Together, these examples illustrate both the promise and the complexity of evaluating AI as it increasingly acts as a collaborator in scientific communication.

Bio: Iryna Gurevych is Professor of Ubiquitous Knowledge Processing in the Department of Computer Science at the Technical University of Darmstadt in Germany. She also is an adjunct professor at MBZUAI in Abu-Dhabi, UAE, and an affiliated professor at INSAIT in Sofia, Bulgaria. She is widely known for fundamental contributions to natural language processing (NLP) and machine learning. Professor Gurevych is a past president of the Association for Computational Linguistics (ACL), the leading professional society in NLP. Her many accolades include being a Fellow of the ACL, an ELLIS Fellow, and the recipient of an ERC Advanced Grant. Most recently, she has received the 2025 Milner award of the British Royal Society for her major contributions to NLP and artificial intelligence that combine deep understanding of human language and cognitive faculty with the latest paradigms in machine learning.
Welcoming AI as a New Colleague: How Should We Evaluate AI for Science?Understanding and Improving Compositional Generalization | AI2Molmo 2 | A new standard for open video intelligenceMolmo 2 | Reasoning across documents and imagesHow far have we come in giving our NLU systems common sense?🚀 Molmo Robotic Demo: AI in ActionOn the Symbiosis of Generative Models and Representation LearningThe BigScience WorkshopThe Pre-trainers toolkit: From dataset construction to model scalingFrom Compression to Convection: A Latent Variable PerspectiveMovement Primitives as Action Sequence Models for Efficient Robot LearningRobot learning and perception for contact-rich manipulation
Ai2 |

Welcoming AI as a New Colleague: How Should We Evaluate AI for Science?

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER