arXiv · 2604.07486
Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
Abstract
Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility. We propose Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and integrates privacy-preserving strategies, including a formal differential privacy (DP) mechanism in the candidate selection, to generate realistic synthetic data. Comprehensive experiments against state-of-the-art private synthetic data generation methods demonstrate that RPSG achieves high fidelity to private data while providing strong privacy protection.
Explore related subjects
Keep this discovery
Qian Ma, Sarah Rajtmajer. 2026-04-08. Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation. https://doi.org/10.18653/v1/2026.findings-acl.10
Cite the original work for its findings. Save a collection to share your selection of sources.