NSDI 26 - Agentix: An Efficient Serving Engine for LLM Agents as General Programs @UsenixOrg
NSDI 26 - Agentix: An Efficient Serving Engine for LLM Agents as General Programs  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Agentix: An Efficient Serving Engine for LLM Agents as General Programs

Michael Luo, University of California, Berkeley, and Google DeepMind; Xiaoxiang Shi, Shanghai Jiao Tong University; Colin Cai, Tianjun Zhang, Justin Wong, and Yichuan Wang, University of California, Berkeley; Chi Wang, Yanping Huang, and Zhifeng Chen, Google DeepMind; Joseph E. Gonzalez and Ion Stoica, University of California, Berkeley

Large language model (LLM) applications are evolving beyond simple chatbots into dynamic, general-purpose agentic programs, which scale LLM calls and output tokens to help AI agents reason, explore, and solve complex tasks. However, existing LLM serving systems ignore dependencies between programs and calls, missing significant opportunities for optimization. Our analysis reveals that programs submitted to LLM serving engines experience long cumulative wait times, primarily due to head-of-line blocking at both the individual LLM request and the program. To address this, we introduce Agentix, an LLM serving system that treats programs as first-class citizens to minimize their end-to-end latencies. Agentix intercepts LLM calls submitted by programs, enriching schedulers with program-level context. We propose two scheduling algorithms—for single-threaded and distributed programs—that preempt and prioritize LLM calls based on their programs' previously completed calls. Our evaluation demonstrates that across diverse LLMs and agentic workloads, Agentix improves throughput of programs by 4-15× at the same latency compared to state-of-the-art systems, such as vLLM.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Agentix: An Efficient Serving Engine for LLM Agents as General ProgramsSREcon24 Europe/Middle East/Africa - You Depend on Time, This Is How It Works and You Won’t...SREcon26 Americas - Epistemology of Incidents and Problem SolvingNSDI 26 - Harp: Improving VPC Network Availability via Efficient Failure Detection and Rerouting...NSDI 26 - Keynote: The Physics of Thought and the Architecture of IntelligenceSREcon25 Europe/Middle East/Africa - Taming the Cost of Telemetry: How Riot Games Reined In...NSDI 26 - HCDN: Coordinated Stream Scheduling for Cost-Effective Live Video DeliveryNSDI 26 - AVA: Towards Agentic Video Analytics with Vision Language ModelsNSDI 26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsNSDI 26 - REAL: Emulating Control Plane at Simulator’s CostPEPR 26 - Enforcement of Data Protection Laws in Africa: Implications for Privacy EngineersPEPR 26 - Adopting AI in Local Government with Privacy and Equity in Mind: A Case Study of the...
USENIX |

NSDI '26 - Agentix: An Efficient Serving Engine for LLM Agents as General Programs

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER