Rethinking LLM efficiency @allenai
Rethinking LLM efficiency  @allenai
Uploaded October 2025 | Updated September 2026, 3 days ago
Large language model (LLM)-based frameworks have achieved remarkable results on complex retrieval and reasoning tasks. However, many of these systems lack thoughtful design choices that balance both efficiency and accuracy. They often operate in an iterative, online, and isolated manner—overlooking relationships between data sources, opportunities for offline processing, and potential for reusability—which results in suboptimal performance. In this talk, I will present my recent work aimed at addressing these limitations across both retrieval and reasoning tasks. Besides enhancing efficiency, our methods also deliver significant improvements in accuracy.

We start by examining inefficiencies in current retrieval frameworks.
1. In part one, we highlight how agent-based querying methods such as ReAct, which iteratively search for relevant information, can be inefficient and miss crucial insights due to their inability to capture relationships across data sources. To overcome this, we propose alignment-oriented retrieval (peterbaile.github.io/arm/), a method that equips LLMs with a structured module to evaluate and exploit data joinability.
2. In part two, we show that current LLM-based online reranking techniques, while effective, often come with significant computational cost and latency for complex retrieval tasks. To tackle this, we introduce EnrichIndex (peterbaile.github.io/enrichindex/), a strategy that enriches the semantic representations of data items in an offline phase.

Finally, we look at reasoning inefficiencies.
3. In part three, we show that current LLMs typically treat each task in isolation, ignoring past interactions, which leads to repeated and re-executed reasoning. To address this, we introduce log-augmented generation (peterbaile.github.io/lag/), a framework that directly reuses prior computation and reasoning from past logs at inference time.

Peter Chen (peterbaile.github.io/) is a PhD student at MIT CSAIL who is fortunate to work with Michael Cafarella, Michael Stonebraker, Samuel Madden, Dan Roth, and Jacob Andreas. His research aims to blend the strengths of data systems and LLM paradigms. His work has appeared at ACL, EMNLP, VLDB, and OSDI, including an Outstanding Paper award at the ACL KnowledgeLM workshop. He is co-organizing the Table Representation Learning (TRL) workshop at NeurIPS 2025. He is also a Schwarzman College of Computing Future Research Cohort fellow, supported by the Croucher Scholarship and Google PhD fellowship in collaboration with MIT.
Rethinking LLM efficiencyWere bringing training text out in the open, Introducing OLMoTraceSide Effects May Include Homogenization and Overusing ClichesGooAQ: Open Question Answering with Diverse Answer Types | AI2Language AI for RNA Virus and RNA VaccineDoing for our robots what nature did for us | Embodied AI Lecture series at AI2Optimal Transport Posterior Alignment for Cross-lingual Semantic ParsingFiguring out how the world works: causality in a world full of real peopleAutomated Hypothesis Validation with Agentic Sequential FalsificationsOpenWebMath: An Open Dataset of High-Quality Mathematical Web TextFrom F to A on the N.Y. Regents Science Exams: An Overview of the Aristo Project | AI2Meet an AI Using Scholar
Ai2 |

Rethinking LLM efficiency

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER