PEPR 26 - Profile-Then-Simulate: Can LLMs Faithfully Generate Differentially Private Synthetic... @UsenixOrg
PEPR 26 - Profile-Then-Simulate: Can LLMs Faithfully Generate Differentially Private Synthetic...  @UsenixOrg
Uploaded June 2026 | Updated September 2026, 3 weeks ago
Profile-Then-Simulate: Can LLMs Faithfully Generate Differentially Private Synthetic Data?

Nassima Bouzid, Capital One

LLM-based simulators are promising new tools for generating synthetic data, especially under conditions that are challenging for traditional Differential Privacy (DP) methods (i.e., high-dimensional tabular data). By condensing customer attributes and behaviors into compact user profiles under DP, we can seed an LLM with data to generate realistic customer transactions. We tested this "Profile-then-Simulate" approach on financial transaction data using PersonaLedger, an LLM-based generator, and compared it to direct DP synthesis on the same dataset.
We found that the LLM-based approach produces usable synthetic data, but direct synthesis still significantly outperforms it on both fraud detection utility and distributional fidelity. We identified systematic LLM biases, not DP noise, as the dominant source of error. The model's learned priors about "typical" financial behavior consistently overrode the statistical distributions we provided as input, particularly for demographic and categorical features, resulting in divergent output data.
This talk shares practical lessons for privacy engineers considering generative AI for synthetic data: (1) LLM biases may dominate DP noise as the primary source of distributional error; (2) direct DP synthesis remains competitive for tractable datasets; and (3) rigorous fidelity evaluation is essential before deploying LLM-generated synthetic data in production pipelines.
Coauthors: Dehao Yuan, Nam H. Nguyen, Mayana Pereira

View the full PEPR '26 program at usenix.org/conference/pepr26/program
PEPR 26 - Profile-Then-Simulate: Can LLMs Faithfully Generate Differentially Private Synthetic...NSDI 26 - ServeGen: Workload Characterization and Generation of Large Language...USENIX ATC 24 - Harmonizing Efficiency and Practicability: Optimizing Resource Utilization in...USENIX Security 25 - ECC.fail: Mounting Rowhammer Attacks on DDR4 Servers with ECC MemoryNSDI 26 - DDoS Detection at the Scale of One Hundred TbpsNSDI 26 - HeteCCL: Synthesizing Near-Optimal Collective Communication Schedules for HeterogeneousNSDI 26 - ZooRoute: Enhancing Cloud-Scale Network Reliability via Candidate Path Provisioning...NSDI 26 - Iris: Expressive Traffic Analysis for the Modern InternetNSDI 26 - CStar Gateway: Augmenting Public Cloud Infrastructure for Heterogeneous Network...NSDI 26 - Heuristic Analysis from Source Code via Symbolic-Guided OptimizationNSDI 26 - Di-PS: System-Algorithm Co-Design for Asynchronous and Heterogeneous Cross-cluster...NSDI 26 - DistVS: Large-scale Vector Search with Compute-Memory Disaggregation
USENIX |

PEPR '26 - Profile-Then-Simulate: Can LLMs Faithfully Generate Differentially Private Synthetic...

SHARE TO X SHARE TO REDDIT SHARE TO FACEBOOK WALLPAPER