Uploaded June 2026 | Updated September 2026, 3 weeks ago
DPSynth: From Research to Production—Engineering Differentially Private Synthetic Tabular Data at Scale
Mikhail Pravilov, Google
Differentially Private (DP) synthetic data is a promising solution for enabling data-driven innovation while protecting user privacy. However, transforming cutting-edge research in DP into robust, scalable, and usable production systems presents significant engineering challenges. Our library, DPSynth, is based on state-of-the-art marginal-based mechanisms (McKenna et al., 2022), and builds upon the foundations of PipelineDP and mbi libraries.
This talk will share our experience in building and applying DPSynth in production settings, highlighting the journey of productionalizing these research concepts. We'll discuss how DPSynth is built to scale for massive datasets using technologies like Apache Beam and Apache Spark. We will also cover key engineering aspects such as handling real-world data constraints to ensure synthetic data utility and validity, and designing for usability with reasonable defaults for non-DP experts. The library is slated for open-source release prior to the conference, aiming to foster wider adoption of practical DP synthetic data techniques.
Authors: Ryan McKenna, Peter Kairouz, Alexander Knop, Vadym Doroshenko, Eva Bertels
View the full PEPR '26 program at usenix.org/conference/pepr26/program
DPSynth: From Research to Production—Engineering Differentially Private Synthetic Tabular Data at Scale
Mikhail Pravilov, Google
Differentially Private (DP) synthetic data is a promising solution for enabling data-driven innovation while protecting user privacy. However, transforming cutting-edge research in DP into robust, scalable, and usable production systems presents significant engineering challenges. Our library, DPSynth, is based on state-of-the-art marginal-based mechanisms (McKenna et al., 2022), and builds upon the foundations of PipelineDP and mbi libraries.
This talk will share our experience in building and applying DPSynth in production settings, highlighting the journey of productionalizing these research concepts. We'll discuss how DPSynth is built to scale for massive datasets using technologies like Apache Beam and Apache Spark. We will also cover key engineering aspects such as handling real-world data constraints to ensure synthetic data utility and validity, and designing for usability with reasonable defaults for non-DP experts. The library is slated for open-source release prior to the conference, aiming to foster wider adoption of practical DP synthetic data techniques.
Authors: Ryan McKenna, Peter Kairouz, Alexander Knop, Vadym Doroshenko, Eva Bertels
View the full PEPR '26 program at usenix.org/conference/pepr26/program










