Lecture 100: InferenceX Continuous OSS Inference Benchmarking @GPUMODE
Lecture 100: InferenceX Continuous OSS Inference Benchmarking  @GPUMODE
Uploaded March 2026 | Updated September 2026, 1 day ago
InferenceX is an open-source (Apache 2.0) automated benchmark designed to keep pace with the rapidly evolving LLM inference software ecosystem. By enabling continuous benchmarking, it addresses the staleness problem of point-in-time evaluations as frameworks like SGLang, vLLM, and TensorRT-LLM release optimizations days apart.

Speakers: Kimbo Chen, Cam Quilici, Bryan Shan
Lecture 100: InferenceX Continuous OSS Inference BenchmarkingLecture 88: TinyTPUcuDNN[Live] ScaleML Series Day 4 — Positional Encodings and PaTH AttentionLecture 4 Compute and Memory BasicsLecture 112: Production Megakernels for Real-World InferenceLecture 18: Fusing KernelsNeighborhood AttentionSmall RL Models at the Speed of Light with LeanRLLecture 53: torch.compile Q&ALecture 11: SparsityLecture 109: TIRx
GPU MODE |

Lecture 100: InferenceX Continuous OSS Inference Benchmarking

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER