InferenceX: Continuous OSS Inference Benchmarking @GPUMODE
InferenceX: Continuous OSS Inference Benchmarking  @GPUMODE
Uploaded March 2026 | Updated September 2026, 10 hours ago
InferenceX is an open-source (Apache 2.0) automated benchmark designed to keep pace with the rapidly evolving LLM inference software ecosystem. By enabling continuous benchmarking, it addresses the staleness problem of point-in-time evaluations as frameworks like SGLang, vLLM, and TensorRT-LLM release optimizations days apart.

Speakers: Kimbo Chen, Cam Quilici, Bryan Shan
InferenceX: Continuous OSS Inference Benchmarking[Live] ScaleML Series Day 2 — Efficient & Effective Long-Context Modeling for Large Language ModelsLecture 95: Single controller programming with MonarchLecture 92: Smol Training PlaybookLecture 49: Low Bit Metal KernelsLive PCCL Fault tolerant collectivesLecture 74: [ScaleML Series] Positional Encodings and PaTH AttentionDistributed ML on consumer devicesLecture 60: Optimizing Linear AttentionLecture 8: CUDA Performance ChecklistLecture 16: On Hands ProfilingLecture 17: NCCL
GPU MODE |

InferenceX: Continuous OSS Inference Benchmarking

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER