NSDI 26 - Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via @UsenixOrg
NSDI 26 - Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via Semantic-Aware Knowledge Caching

Chaoyi Ruan, National University of Singapore; Chao Bi, University of Science and Technology of China; Kaiwen Zheng, University of Toronto; Ziji Shi, National University of Singapore; Xinyi Wan, Sea AI Lab and National University of Singapore; Jialin Li, National University of Singapore

Large Language Model (LLM) agents tackle data-intensive tasks such as deep research and code generation. However, their effectiveness depends on frequent interactions with knowledge sources across remote clouds or regions. Such interactions can create non-trivial latency and cost bottlenecks. Existing caching solutions focus on exact-match queries, limiting their effectiveness for semantic knowledge reuse.

To address this challenge, we introduce Cortex, a novel cross-region knowledge caching architecture for LLM agents. At its core are two abstractions: Semantic Element (SE) and Semantic Retrieval Index (Seri). A semantic element captures the semantic embedding representation of an LLM query together with performance-aware metadata such as latency, cost, and staticity. Seri then provides two-stage retrieval: a vector similar index with semantic embedding for fast candidate selection and a lightweight LLM-powered semantic judger for precise validation. Atop these primitives, Cortex builds a new cache interface that includes a new semantic-aware cache hit definition, a cost-efficient eviction policy, and proactive prefetching. To reduce overhead, Cortex co-locates the small LLM judger with the main LLM using adaptive scheduling and resource sharing. Our evaluation demonstrates that Cortex delivers substantial performance improvements without compromising correctness. On representative search workloads, Cortex achieves up to a 3.6× increase in throughput by maintaining cache hit rates of over 85×, while preserving accuracy virtually identical to non-cached baselines. Cortex also improves throughput for coding tasks by 20×, showcasing its versatility across diverse agentic workloads.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM viaNSDI 26 - Offloading Cloud Network Services at Production Scale with SONiC DASH SmartSwitchNSDI 26 - Net-P4ct: Enhanced WAN Bandwidth Fair Sharing Using P4 Programmable SwitchesNSDI 26 - MAE: More Adaptive Video Encoder for Consistent Low Latency in High-Quality Real-Time...USENIX ATC 24 - ScalaAFA: Constructing User-Space All-Flash Array Engine with Holistic DesignsNSDI 26 - EZ-SAVE: Evaluation of Easy-to-Deploy Source Address Validation PoliciesNSDI 26- Bridging Storage and Execution: A Semantic Virtual Bus for On-Demand Application StreamingPEPR 26 - Mapping the Privacy Workforce in the AI EraNSDI 26 - EROICA: Online Performance Troubleshooting for Large-scale Model TrainingNSDI 26 - CrossCheck: Input Validation for WAN Control SystemsNSDI 26 - Controlling Arbitrary Internet Queues with TitrateNSDI 26 - Learning to Tune Optical WANs: A Field Deployment of Noise Models in Optical Networks
USENIX |

NSDI '26 - Cortex: Achieving Low-Latency, Cost-Efficient Remote Data Access For LLM via

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER