NSDI 26 - Phantora: Maximizing Code Reuse in Simulation-based Machine Learning... @UsenixOrg
NSDI 26 - Phantora: Maximizing Code Reuse in Simulation-based Machine Learning...  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
NSDI '26 - Phantora: Maximizing Code Reuse in Simulation-based Machine Learning System Performance Estimation

Jianxing Qin, Duke University; Jingrong Chen, Uber; Xinhao Kong, NVIDIA; Yongji Wu, University of California, Berkeley; Tianjun Yuan, Duke University; Liang Luo, Zhaodong Wang, and Ying Zhang, Meta; Tingjun Chen, Alvin R. Lebeck, and Danyang Zhuo, Duke University

Modern machine learning (ML) training workloads place substantial demands on both computational and communication resources. Consequently, accurate performance estimation has become increasingly critical for guiding system design decisions, such as the selection of parallelization strategies, cluster configurations, and hardware provisioning. Existing simulation-based performance estimation requires reimplementing the ML framework in a simulator, which demands significant manual effort and is hard to maintain as ML frameworks evolve rapidly.

This paper introduces Phantora, a hybrid GPU cluster simulator designed for performance estimation of ML training workloads. Phantora executes unmodified ML frameworks as is within a distributed, containerized environment. Each container emulates the behavior of a GPU server in a large-scale cluster, while Phantora intercepts and simulates GPU- and communication-related operations to provide high-fidelity performance estimation. We call this approach hybrid simulation of ML systems, in contrast to traditional methods that simulate static workloads. The primary advantage of hybrid simulation is that it allows direct reuse of ML framework source code in simulation, avoiding the need for reimplementation. Our evaluation shows that Phantora provides accuracy comparable to static workload simulation while supporting three state-of-the-art LLM training frameworks out-of-the-box. In addition, Phantora operates on a single GPU, eliminating the need for the resource-intensive trace collection and workload extraction steps required by traditional trace-based simulators. Phantora is open-sourced at github.com/QDelta/Phantora.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Phantora: Maximizing Code Reuse in Simulation-based Machine Learning...USENIX Security 24 - Practical Data-Only Attack GenerationNSDI 26 - Starfish: A Topology-Routing Co-Design for Small-Scale Data CentersSREcon26 Americas - How We Debug 1000s of Databases with AI: Lessons from an AI-Assisted Database...NSDI 26 - SmartNIC-Enabled Live Migration for Storage-Optimized VMs with PYROCUMULUSSREcon26 Americas - The Gashlycrumb Tinies of AI Networking You Must Know (or Languish!)NSDI 25 - CATO: End-to-End Optimization of ML-Based Traffic Analysis PipelinesNSDI 26 - BBC: Enabling BLE to Support Bluetooth ClassicUSENIX ATC 24 - Fast Inference for Probabilistic Graphical ModelsNSDI 26 - FLARE: Anomaly Diagnostics for Divergent LLM Training in GPU Clusters of Thousand-Plus...NSDI 26 - Remote TCP Connection Offload and ApplicationsUSENIX Security 25 - GraphAce: Secure Two-Party Graph Analysis Achieving Communication Efficiency
USENIX |

NSDI '26 - Phantora: Maximizing Code Reuse in Simulation-based Machine Learning...

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER