Uploaded June 2026 | Updated September 2026, 3 weeks ago
Fractal: Fault-Tolerant Shell-Script Distribution
Zhicheng Huang, Ramiz Dundar, and Yizheng Xie, Brown University; Konstantinos Kallas, University of California, Los Angeles; Nikos Vasilakis, Brown University
This paper presents FRACTAL, a new system that offers fault tolerant distributed shell script execution for unmodified scripts. FRACTAL first distinguishes recoverable regions from side-effectful ones, and augments them with additional runtime support aimed at fault recovery. It employs precise dependency and progress tracking at the subgraph level to offer sound and efficient fault recovery. It minimizes the number of upstream regions that are re-executed during recovery and ensures exactly-once semantics upon recovery for downstream regions. Evaluation on 4- and 30-node clusters indicates average fault-free speedups of (1) greater than 9.6x over Bash, a single-node shell-interpreter baseline, (2) greater than 5.5x over Hadoop Streaming, a MapReduce system that supports language-agnostic third-party components, and (3) 17% over DiSh, a state-of-the-art fault-intolerant shell-script distribution system—all while recovering 7.8–16.4x faster than Hadoop Streaming in cases of faults.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
Fractal: Fault-Tolerant Shell-Script Distribution
Zhicheng Huang, Ramiz Dundar, and Yizheng Xie, Brown University; Konstantinos Kallas, University of California, Los Angeles; Nikos Vasilakis, Brown University
This paper presents FRACTAL, a new system that offers fault tolerant distributed shell script execution for unmodified scripts. FRACTAL first distinguishes recoverable regions from side-effectful ones, and augments them with additional runtime support aimed at fault recovery. It employs precise dependency and progress tracking at the subgraph level to offer sound and efficient fault recovery. It minimizes the number of upstream regions that are re-executed during recovery and ensures exactly-once semantics upon recovery for downstream regions. Evaluation on 4- and 30-node clusters indicates average fault-free speedups of (1) greater than 9.6x over Bash, a single-node shell-interpreter baseline, (2) greater than 5.5x over Hadoop Streaming, a MapReduce system that supports language-agnostic third-party components, and (3) 17% over DiSh, a state-of-the-art fault-intolerant shell-script distribution system—all while recovering 7.8–16.4x faster than Hadoop Streaming in cases of faults.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions










