Uploaded March 2026 | Updated September 2026, 10 hours ago
InferenceX is an open-source (Apache 2.0) automated benchmark designed to keep pace with the rapidly evolving LLM inference software ecosystem. By enabling continuous benchmarking, it addresses the staleness problem of point-in-time evaluations as frameworks like SGLang, vLLM, and TensorRT-LLM release optimizations days apart.
Speakers: Kimbo Chen, Cam Quilici, Bryan Shan
InferenceX is an open-source (Apache 2.0) automated benchmark designed to keep pace with the rapidly evolving LLM inference software ecosystem. By enabling continuous benchmarking, it addresses the staleness problem of point-in-time evaluations as frameworks like SGLang, vLLM, and TensorRT-LLM release optimizations days apart.
Speakers: Kimbo Chen, Cam Quilici, Bryan Shan
![[Live] ScaleML Series Day 2 — Efficient & Effective Long-Context Modeling for Large Language Models
Day 2: Efficient & Effective Long-Context Modeling for Large Language Models by Guangxuan Xiao.
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. [Live] ScaleML Series Day 2 — Efficient & Effective Long-Context Modeling for Large Language Models](https://i.ytimg.com/vi/PKYvAc9UZhk/mqdefault.jpg)




![Lecture 74: [ScaleML Series] Positional Encodings and PaTH Attention
Speaker: Songlin Yang (PhD at MIT).
Full Schedule: https://scale-ml.org/bootcamp/
The GPU MODE x Scale ML speaker series is a 5-day, online event hosted on the GPU MODE YouTube channel where top researchers in AI will talk about various architectural and system-level advances that are integrated into OpenAI’s frontier open-source model, GPT-OSS.
Each day will consist of ~2 hours of talks and discussions (around noon PST, may start at slightly different times each day so please check frequently), covering a different component of the evolving transformer stack—from quirks in the attention mechanism and positional encodings to quantization, MoEs, and custom GPU kernels. Lecture 74: [ScaleML Series] Positional Encodings and PaTH Attention](https://i.ytimg.com/vi/QXbXdN3KIcY/mqdefault.jpg)




