Are long context LLMs the death of RAG? @ai-science
Are long context LLMs the death of RAG?  @ai-science
Uploaded March 2024 | Updated September 2026, 3 weeks ago
Check out my essays: aisc.substack.com
OR book me to talk: calendly.com/amirfzpr
OR subscribe to our event calendar: https://lu.ma/aisc-llm-school

AF: You said people are saying there's long context LLMs, therefore RAG is not that interesting. I want to go back to that and underline. Why do you disagree?

AM: I think some people misunderstand what RAG is good for when they say that. Gemini with its really long context window, I can throw all the text I want at this and it can reference it inside its own context window. Why do I need RAG? For me, RAG is not just about giving an answer based on that information. I think having a data set that you manage separately is quite important from an audit point of view. In our line of work in financial services, it's important to see what's the whole lineage of the data you've used to come to this answer. Fundamentally, the data is sacrosanct.

I think RAG's going to stay for a long time. With the large context window model, it probably means your RAG systems can become more powerful, or they can do more things at one time, or they can solve slightly more complex problems.

AF: Maybe we need a new term for it because RAG has moved on from how it was proposed originally. At the retriever stage, there are so many things that we are layering in these days, there is the retrieval ranking, there is the privacy controls, there is the PII handling, there is access control, security control, domain knowledge. So many things we're doing there that a long context window is not a solution to.

Even with a 4000 token context window, you're seeing a lot of "lost in the middle" type of problems. When you fill up the context, they're not particularly good at figuring out where to pay attention to exactly.

MA: I think everything you've said about those other elements are super important when it comes to enterprise scale stuff. All of this architecture here, there's a lot of things in here that I will want to stick around for a long time. If I just replace this with "talk to an LLM", I lose so much control. I lose so much auditability. I lose so much security.
Are long context LLMs the death of RAG?Lateral Thinking Example in Deep Research SystemLimitations of Agentic Frameworks: When to Use a Custom FrameworkHow Do State Machines Work?What is the relationship between LLMs and multi-modality?Securing Your LLMs: The OWASP Top Risks You Can’t Ignore5 Commandments of Building LLM ProductsSage Social Studio: AI Application that Polishes Content for LinkedIn, Substack & TwitterSofIA - The Ultimate AI Guide for Canadian ImmigrantsWhat Makes DeepSeek R1 Multi-token Prediction Unique?Evaluating Job Exposure to Large Language ModelsAI Agents & Game Development: Why ChatGPT Isn’t Enough for D&D (And What I Built Instead)
LLMs Explained - Aggregate Intellect - AI.SCIENCE |

Are long context LLMs the death of RAG?

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER