Uploaded June 2026 | Updated September 2026, 3 weeks ago
PD3: Prefetching Data with DPUs for Disaggregated Memory
Sidharth Sankhe, Felix Zhang, and Umayrah Chonee, University of Toronto; Sherman Lim, National University of Singapore; Jiasheng Hu, University of Toronto; Jialin Li, National University of Singapore; Qizhen Zhang, University of Toronto
We introduce PD3, a memory disaggregation solution that "avoids" cache misses, via prefetching, on compute servers and thus all their associated overhead. Unlike a traditional prefetcher that may pollute the cache or miss preloading opportunities due to false positives and false negatives, PD3 prevents mis-predictions with network support and minimal yet critical application information. Enabling PD3 is data processing units or DPUs, which allow (1) parsing user requests before they are processed by the compute server, (2) fetching data from remote memory on the shortest path, (3) offloading expensive RDMA and DMA operations from the host, and (4) incorporating application knowledge to faithfully predict cache misses and take actions accordingly. Designing PD3 requires reconciling DPU resource constraints and scaling requirements of cloud data systems, as well as achieving high efficiency with a myriad of performance optimizations. Our experimental results on real hardware, applications, and workloads show that with nominal compute-local memory, PD3 eliminates the performance gap between memory-disaggregated applications and their monolithic counterparts.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions
PD3: Prefetching Data with DPUs for Disaggregated Memory
Sidharth Sankhe, Felix Zhang, and Umayrah Chonee, University of Toronto; Sherman Lim, National University of Singapore; Jiasheng Hu, University of Toronto; Jialin Li, National University of Singapore; Qizhen Zhang, University of Toronto
We introduce PD3, a memory disaggregation solution that "avoids" cache misses, via prefetching, on compute servers and thus all their associated overhead. Unlike a traditional prefetcher that may pollute the cache or miss preloading opportunities due to false positives and false negatives, PD3 prevents mis-predictions with network support and minimal yet critical application information. Enabling PD3 is data processing units or DPUs, which allow (1) parsing user requests before they are processed by the compute server, (2) fetching data from remote memory on the shortest path, (3) offloading expensive RDMA and DMA operations from the host, and (4) incorporating application knowledge to faithfully predict cache misses and take actions accordingly. Designing PD3 requires reconciling DPU resource constraints and scaling requirements of cloud data systems, as well as achieving high efficiency with a myriad of performance optimizations. Our experimental results on real hardware, applications, and workloads show that with nominal compute-local memory, PD3 eliminates the performance gap between memory-disaggregated applications and their monolithic counterparts.
View the full NSDI '26 program at usenix.org/conference/nsdi26/technical-sessions










