NSDI 26 - AVA: Towards Agentic Video Analytics with Vision Language Models @UsenixOrg
NSDI 26 - AVA: Towards Agentic Video Analytics with Vision Language Models  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
NSDI '26 - AVA: Towards Agentic Video Analytics with Vision Language Models

Yuxuan Yan, Zhejiang University; Shiqi Jiang, Microsoft Research; Ting Cao, Tsinghua University; Yifan Yang, Microsoft Research; Qianqian Yang and Yuanchao Shu, Zhejiang University; Yuqing Yang and Lili Qiu, Microsoft Research

AI-driven video analytics has become increasingly important across diverse domains. However, existing systems are often constrained to specific, predefined tasks, limiting their adaptability in open-ended analytical scenarios. The recent emergence of Vision Language Models (VLMs) as transformative technologies offers significant potential for enabling open-ended video understanding, reasoning, and analytics. Nevertheless, their limited context windows present challenges when processing ultra-long video content, which is prevalent in real-world applications. To address this, we introduce AVA, a VLM-powered system designed for open-ended, advanced video analytics. AVA incorporates two key innovations: (1) the near real-time construction of Event Knowledge Graphs (EKGs) for efficient indexing of long or continuous video streams, and (2) an agentic retrieval-generation mechanism that leverages EKGs to handle complex and diverse queries. Comprehensive evaluations on public benchmarks, LVBench and VideoMME-Long, demonstrate that AVA achieves state-of-the-art performance, attaining 62.3% and 64.1% accuracy, respectively, significantly surpassing existing VLM and video Retrieval-Augmented Generation (RAG) systems. Furthermore, to evaluate video analytics in ultra-long and open-world video scenarios, we introduce a new benchmark, AVA-100. This benchmark comprises 8 videos, each exceeding 10 hours in duration, along with 120 manually annotated, diverse, and complex question-answer pairs. On AVA-100, AVA achieves top-tier performance with an accuracy of 75.8%.

The source code of AVA is available at github.com/I-ESC/Project-Ava. The AVA-100 benchmark could be accessed at huggingface.co/datasets/iesc/Ava-100.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - AVA: Towards Agentic Video Analytics with Vision Language ModelsNSDI 26 - DroidSpeak: KV Cache Sharing Across Fine-tuned Model VariantsNSDI 26 - REAL: Emulating Control Plane at Simulator’s CostPEPR 26 - Enforcement of Data Protection Laws in Africa: Implications for Privacy EngineersPEPR 26 - Adopting AI in Local Government with Privacy and Equity in Mind: A Case Study of the...SREcon26 Americas - Stop Reading Changelogs: Safer Kubernetes Upgrades with SimulationNSDI 26 - Skyline: A Cloud Centric Internet Monitoring EngineNSDI 26 - HybridMesh: A Hardware-software Hybrid Approach for Accelerating Service Mesh IngressSREcon26 Americas - AI Agents for Incident Investigation: The Good, The Bad, and The UglySREcon26 Americas - 5 Wrong Hypotheses about PostgreSQL Multi-Transaction LocksSREcon25 Europe/Middle East/Africa - Utilization Is the Key to Efficiency: What It Takes to Run...NSDI 26 - PolicyCache: Intra-flow Learning in Congestion Control
USENIX |

NSDI '26 - AVA: Towards Agentic Video Analytics with Vision Language Models

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER