When Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric Memories @allenai
When Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric Memories  @allenai
Uploaded June 2023 | Updated September 2026, 1 day ago
Presentation of ACL 2023 main conference long paper "When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories".

Alex Mallen*, Akari Asai*, Victor Zhong, Rajarshi Das, Daniel Khashabi, Hannaneh Hajishirzi

Despite their impressive performance on diverse tasks, large language models (LMs) still struggle with tasks requiring rich world knowledge, implying the limitations of relying solely on their parameters to encode a wealth of world knowledge. This paper aims to understand LMs' strengths and limitations in memorizing factual knowledge, by conducting large-scale knowledge probing experiments of 10 models and 4 augmentation methods on PopQA, our new open-domain QA dataset with 14k questions. We find that LMs struggle with less popular factual knowledge, and that scaling fails to appreciably improve memorization of factual knowledge in the long tail. We then show that retrieval-augmented LMs largely outperform orders of magnitude larger LMs, while unassisted LMs remain competitive in questions about high-popularity entities. Based on those findings, we devise a simple, yet effective, method for powerful and efficient retrieval-augmented LMs, which retrieves non-parametric memories only when necessary. Experimental results show that this significantly improves models' performance while reducing the inference costs.

Camera ready paper: arxiv.org/abs/2212.10511
Code: github.com/AlexTMallen/adaptive-retrieval
When Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric MemoriesOpen-Ended Learning Leads to Generally Capable Agents | Embodied AI Lecture Series at AI2From LLMs to Agents: Generalizability from the Inside OutData-Centric Approaches to Adapting Foundation ModelsThe University of Washington eScience Institute: a Home for Data-Intensive DiscoveryGeneralization for Robot Learning In The Wild | Embodied AI Lecture series at AI2Towards Generalist Agents for Accelerating Scientific Discovery171Modular Language ModelsTowards robust long-form text generation systemsDigital Socrates: Evaluating LLMs through Explanation CritiquesAi2: The best AI tools arent built for scientists. Theyre built with them.Beyond End-to-end: Decomposed Modeling and Representations in NLP
Ai2 |

When Not to Trust Language Models: Investigating Effectiveness of Parametric&Non-Parametric Memories

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER