[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84478-en":3,"doc-seo-84478-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84478,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation","Large language models are used to generate synthetic data, especially when replicating private text. Producing such replicas requires careful balancing of privacy and utility. This paper proposes Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and adds privacy-preserving strategies, including a formal differential privacy mechanism during candidate selection. Experiments against state-of-the-art private synthetic data methods show high fidelity to private data with strong privacy protection.","Private Seeds, Public LLMs: Realistic and Privacy-Preserving  \nSynthetic Data Generation  \nQian Ma, Sarah Rajtmajer  \nInformation Sciences and Technology, The Pennsylvania State University {qfm5033, [smr48}@psu.edu](smr48}@psu.edu)  \narXiv :2604 .07486v 3 [ cs .CR] 13 Jul 2026  \nAbstract  \nLarge language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing synthetic replicas of private text, which requires carefully balancing privacy and utility.  \nWe propose Realistic and Privacy-Preserving Synthetic Data Generation (RPSG), which uses private seeds and integrates privacy-preserving strategies, including a formal differential privacy (DP) mechanism in the candidate selection, to generate realistic synthetic data. Comprehensive experiments against state-of-theart private synthetic data generation methods demonstrate that RPSG achieves high fidelity to private data while providing strong privacy protection. 1  \n1 Introduction  \nSynthetic data generation is an active area of work in natural language processing (NLP) (Bommasaniet al., 2019 ; Yu et al., 2022) with applications ranging from clinical text analysis (Walonoski et al., 2018 ; Tang et al., 2023) to social media synthesis (Cao et al., 2023 ; Lu et al., 2023) .  \nOne type of synthetic data of significant practical interest is synthetic replicas of private text (Hou et al., 2024 ; Yu et al., 2023) . Oftentimes, text data that would be of benefit, e.g., to researchers, policymakers, or technologists, cannot be shared due to privacy considerations. Text collected from social media platforms, for example, often contains users’ voluntarily disclosed personal information–so-called self-disclosures (see (Ashuri and Halperin, 2024) for a recent interdisciplinary review) . Such text can be used for user targeting and manipulation; if the personal information shared is more sensitive, e.g., personally identifiable information (PII), the risks can be graver (Gruzd and Hernández-García, 2018) .  \n1All code is available at [https://github.com/masonmq/](https://github.com/masonmq/)[ ](https://github.com/masonmq/)rpsg  \nMainstream methods for generating privacypreserving synthetic data include fine-tuning LLMs (e.g., DistilGPT2 (Wolf et al., 2020)) with differential privacy (DP) mechanisms applied through extensive modifications of gradient descent during training (Hou et al., 2024), as well as using prompt engineering (e.g., GPT-4 (OpenAI, 2023)) to guide models toward producing semantically similar synthetic data (Yukhymenko et al., 2024) . However, fine-tuning approaches without rigorous privacy safeguards have been shown to suffer from memorization and leakage vulnerabilities (Mireshghallah et al., 2022 ; Li et al., 2024a), and are increasingly impractical as many modern LLMs are accessible only via APIs. In contrast, prompt-based methods are model-access agnostic and can be applied to both API-based and open-source models. However, they struggle to generate high-quality synthetic data that preserves the richness and utility of the original private data (Xie et al., 2024) .  \nIn response to these limitations, we propose a novel method for realistic and privacy-preserving synthetic data generation (RPSG) (Figure 1) . RPSG is designed using private data as seeds to generate high-quality synthetic data that closely resembles the original while reducing privacy risk. In Phase 1, an abstraction model produces multiple sentiment-aligned abstracted candidates for each private seed, reducing identifiable patterns and semantic structures; a formal DP mechanism is then applied to select candidates. In Phase 2, an LLM generates variations of these DP-selected candidates. In Phase 3, synthetic variants are refined to reduce memorization risks and redact any PII. The resulting samples constitute a final set of realistic and privacy-preserving synthetic data, suitable for downstream applications.  \nWe conduct comprehensive expe","cbCaicuub7sM5BC7","https://ap.wps.com/l/cbCaicuub7sM5BC7","pdf",685227,1,22,"English","en",105,"# Introduction\n## Synthetic replicas of private text\n## Limitations of existing privacy-preserving methods\n## The RPSG method pipeline\n## Experimental evaluation and contributions","[{\"question\":\"What problem does RPSG address in synthetic data generation?\",\"answer\":\"RPSG targets generating realistic synthetic replicas of private text while balancing privacy risk and data utility for downstream use.\"},{\"question\":\"How does RPSG incorporate privacy protection?\",\"answer\":\"RPSG applies a formal differential privacy mechanism during candidate selection in Phase 1, and further refines outputs to reduce memorization risks and redact any PII.\"},{\"question\":\"How is RPSG evaluated and what does it achieve compared with baselines?\",\"answer\":\"Experiments use a PubMed benchmark dataset and an original Reddit dataset, comparing against gradient-based and prompt-based DP baselines plus a one-to-one rewriting baseline. Results report improved fidelity, diversity, lexical quality, and stronger resistance to membership inference attacks.\"}]",1784195928,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"private-seeds-public-llms-realistic-and-privacy-preserving-synthetic-data-generation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/private-seeds-public-llms-realistic-and-privacy-preserving-synthetic-data-generation/84478/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RPSG address in synthetic data generation?","Question",{"text":75,"@type":76},"RPSG targets generating realistic synthetic replicas of private text while balancing privacy risk and data utility for downstream use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RPSG incorporate privacy protection?",{"text":80,"@type":76},"RPSG applies a formal differential privacy mechanism during candidate selection in Phase 1, and further refines outputs to reduce memorization risks and redact any PII.",{"name":82,"@type":73,"acceptedAnswer":83},"How is RPSG evaluated and what does it achieve compared with baselines?",{"text":84,"@type":76},"Experiments use a PubMed benchmark dataset and an original Reddit dataset, comparing against gradient-based and prompt-based DP baselines plus a one-to-one rewriting baseline. Results report improved fidelity, diversity, lexical quality, and stronger resistance to membership inference attacks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]