NSDI 26 - Detecting and Diagnosing Errors in Serving Archived Web Pages @UsenixOrg
NSDI 26 - Detecting and Diagnosing Errors in Serving Archived Web Pages  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Detecting and Diagnosing Errors in Serving Archived Web Pages

Jingyuan Zhu, University of Michigan; Huanchen Sun and Harsha V. Madhyastha, University of Southern California
Community Award Winner!

Web archives crawl and save copies of pages from the web, enabling users to interact with web pages in the form they existed in the past. Prior to serving any archived page, an archive rewrites the page’s source so that users’ browsers fetch the page’s resources from the archive, not from the servers which originally hosted them. But, on many modern pages, an archive’s edits to crawled scripts result in a loss of fidelity, i.e., an archived copy fails to accurately mimic the original page even when the archive had crawled all resources on the page.

To help the developers of archival systems identify and fix the bugs which result in incorrect rewrites of crawled pages, we present FidEx. First, FidEx enables accurate identification of the pages on which an archive violates fidelity. It does so by tracking and comparing the execution of scripts between when a page is crawled and when its copy is loaded. In comparison to existing methods which compare the two loads using either screenshots or the errors reported by the browser, FidEx reduces the false positive rate from around 70% to less than 10%. Second, on every page on which it identifies a loss of fidelity, FidEx pinpoints which subset of the archive’s edits to the page are erroneous. Leveraging this input to fix bugs in the most widely used archival system, we reduced the fraction of archived pages which violate fidelity from 15% to 9%.

View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
NSDI 26 - Detecting and Diagnosing Errors in Serving Archived Web PagesNSDI 26 - SYMI: Efficient Mixture-of-Experts Training via Model and Optimizer State DecouplingPEPR 26 - Vision: Human-as-the-Unit Privacy Management with AI AgentsPEPR 26 - Surfacing Hidden Privacy Risks in Code: Lessons from LLM and Retrieval Assisted DetectionNSDI 26 - The GOODPUT System: A Machine Learning-Driven Optimization Framework for Dynamic...NSDI 26 - PrvTel: Lightweight Models for Private and Accurate Telemetry Data RetentionNSDI 26 - Count-Based Abstractions for Performance Verification of Contention PointsNSDI 26 - A Systematic Threat Analysis and Practical Attacks on Automated Frequency CoordinationNSDI 26 - ForestColl: Throughput-Optimal Collective Communications on Heterogeneous Network FabricsNSDI 26 - CacheCatalyst: Enhancing Web Caching for the Latency-Constrained InternetNSDI 26 - BURST: Seeking High-performance, Interoperability and Scalability in Soft-RDMANSDI 26 - MirrorNet: High-fidelity and Scalable Network Emulation for Software-defined WAN
USENIX |

NSDI '26 - Detecting and Diagnosing Errors in Serving Archived Web Pages

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER