[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84889-en":3,"doc-seo-84889-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84889,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems Reflections from the Breast Cancer Domain","Modern software increasingly relies on data for analysis, prediction, testing, and decision-making, yet domains like medicine, safety-critical systems, and regulated industries often lack abundant, shareable, representative datasets. Synthetic data generation is proposed as a remedy, but experience engineering intraoperative radiotherapy software for breast cancer indicates synthetic data shifts rather than removes the core engineering difficulty. The work frames the problem as property-driven synthetic data engineering, focusing on which validity properties to preserve, how to elicit them from stakeholders, validate them under privacy constraints, and evolve them across pipelines.","Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain  \nAurora Francesca Zanenga  \n[aurora.zanenga@unibg.it](aurora.zanenga@unibg.it)[ ](aurora.zanenga@unibg.it)University of Bergamo Bergamo, Italy  \nSaverio D’Amico  \nsaverio.damico@humanitas.it Humanitas Clinical and Research Center, IRCCSRozzano, Italy  \nAndrea Bombarda  \n[andrea.bombarda@unibg.it](andrea.bombarda@unibg.it)[ ](andrea.bombarda@unibg.it)University of Bergamo Bergamo, Italy  \nRita De Sanctis  \n[rdesanctis@asst-pg23.it](rdesanctis@asst-pg23.it)[ ](rdesanctis@asst-pg23.it)ASST Papa Giovanni XXIII Bergamo, Italy  \nMarsha Chechik  \n[chechik@cs.toronto.edu](chechik@cs.toronto.edu)[ ](chechik@cs.toronto.edu)University of Toronto Toronto, Canada  \nAlberto Zambelli  \n[azambelli@asst-pg23.it](azambelli@asst-pg23.it)[ ](azambelli@asst-pg23.it)ASST Papa Giovanni XXIII Bergamo, Italy  \narXiv :2607 .06 133v 1 [ cs . SE] 7 Jul 2026  \nClaudio Menghi  \n[claudio.menghi@unibg.it](claudio.menghi@unibg.it)[ ](claudio.menghi@unibg.it)University of Bergamo Bergamo, Italy McMaster University Hamilton, ON, Canada  \nAbstract  \nModern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, including medicine, safety-critical systems, and regulated industries, lack abundant, shareable, or representative data. Synthetic data generation is often proposed as a remedy, but our experience engineering software for intraoperative radiotherapy (IORT) in breast cancer treatment suggests that synthetic data shifts rather than solves the central engineering problem. The key challenge becomes deciding which properties synthetic data must preserve, how these properties should be elicited from stakeholders, how they can be validated under privacy constraints, and how they evolve. We call this problem property-driven synthetic data engineering. Drawing on a collaboration with oncologists and preliminary experiments with a sensitive IORT dataset, we identify challenges in requirements, validation, privacy, and pipeline evolution. We argue that automated software engineering research should develop methods and tools for eliciting, formalizing, checking, and evolving validity properties for synthetic data in data-scarce software systems.  \nCCS Concepts  \n• Software and its engineering;  \nKeywords  \nProperty-Driven Synthetic Data, Breast Cancer, Data-Scarce Software Systems  \n1 Introduction  \nData are driving modern software applications. Tasks such as prediction, classification, testing, and decision-making rely on the availability of large, accessible datasets, and the increasing adoption of artificial intelligence further reinforces this assumption. However,  \nin many safety-critical and privacy-sensitive domains, this assumption does not hold. In domains such as medicine and regulated industries, data are often scarce, difficult to access, and subject to strict confidentiality constraints. As a result, software engineering (SE) practices that implicitly rely on abundant data become difficult to apply. This is the case with our partner, a large hospital with more than one million outpatient services and 100,000 emergency department visits. Despite the large number of patients, in many medical scenarios, the number of patients affected by certain pathologies is (fortunately) limited, and the data contain sensitive health information. For example, in the case of breast cancer, while the overall number of cases is significant (approximately 2.3 million new cases each year [13]), the therapeutic landscape [5] is rapidly evolving, and the number of patients within specific subpopulations (e.g., young women or minority groups) may be very small or even absent. Consequently, engineering software relying on such data introduces specific challenges that contrast with mainstream SE practices based on abundant and shareable datasets.  \nTo address data scarcity, Synthetic Data Ge","cbCaivtkcGzovVTj","https://ap.wps.com/l/cbCaivtkcGzovVTj","pdf",462847,2,1,5,"English","en",105,"# Introduction\n## Data scarcity in safety-critical and privacy-sensitive domains\n## Synthetic data generation as a shifted challenge\n## Property-driven synthetic data engineering proposal","[{\"question\":\"Why does data scarcity break standard software engineering assumptions?\",\"answer\":\"In safety-critical and privacy-sensitive domains such as medicine, datasets are scarce, access is restricted, and data may be sensitive and non-shareable. This makes practices that rely on abundant ground-truth data difficult to apply.\"},{\"question\":\"What does the paper claim about synthetic data generation?\",\"answer\":\"Synthetic data does not eliminate the central problem; it shifts it. The key difficulty becomes choosing which properties must be preserved and how to elicit, validate, and maintain them.\"},{\"question\":\"What is property-driven synthetic data engineering?\",\"answer\":\"It is a research agenda where stakeholder-specific validity properties are treated as the central artifacts. These properties must be elicited, formalized, checked, and evolved under changing requirements and privacy constraints.\"}]",1784199040,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"property-driven-synthetic-data-engineering-for-data-scarce-software-systems-reflections-from-the-breast-cancer-domain","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/property-driven-synthetic-data-engineering-for-data-scarce-software-systems-reflections-from-the-breast-cancer-domain/84889/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does data scarcity break standard software engineering assumptions?","Question",{"text":75,"@type":76},"In safety-critical and privacy-sensitive domains such as medicine, datasets are scarce, access is restricted, and data may be sensitive and non-shareable. This makes practices that rely on abundant ground-truth data difficult to apply.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper claim about synthetic data generation?",{"text":80,"@type":76},"Synthetic data does not eliminate the central problem; it shifts it. The key difficulty becomes choosing which properties must be preserved and how to elicit, validate, and maintain them.",{"name":82,"@type":73,"acceptedAnswer":83},"What is property-driven synthetic data engineering?",{"text":84,"@type":76},"It is a research agenda where stakeholder-specific validity properties are treated as the central artifacts. These properties must be elicited, formalized, checked, and evolved under changing requirements and privacy constraints.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]