[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83454-en":3,"doc-seo-83454-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83454,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","A Filtered Mixture-of-Generators for Fully Synthetic Survival Training","Survival analysis models time-to-event outcomes in critical domains such as oncology and cardiology, but real training data are costly to collect, cohorts are typically small, and privacy rules limit cross-institution sharing. Tabular generative models could enable fully synthetic training, yet single generators are often data-hungry and fail to match real-data performance. FoGS reframes synthetic construction as sample selection from a heterogeneous generator pool, using survival-model ensembles and an outer-loop optimization to maximize downstream C-index and IBS.","arXiv :2607 .00127v1 [ cs .LG] 30 Jun 2026  \nA Filtered Mixture-of-Generators for Fully Synthetic  \nSurvival Training  \nNiccolò Maria Rizzi 1 , Eugenio Lomurno* 1 , Alberto Archetti 1 , and Matteo Matteucci 1  \n1 Politecnico di Milano, Milan, Italy  \nAbstract  \nObjective: Survival analysis is a statistical framework for time-to-event modelling in a wide range of critical domains. In clinical settings, training data are particularly costly to assemble, since events accrue over years of follow-up, cohort sizes remain small, and privacy regulations restrict sharing across institutions. Tabular generative models offer, in principle, both augmentation and privacy-preserving cohort sharing, but are themselves data-hungry: on the small cohorts typical of  \nsurvival analysis, a single generator rarely characterizes the population well enough for downstream models trained on its output to match real-data performance. We aim to make fully synthetic training a viable substitute for real-data training in this regime.  \nMethods: We propose FoGS (Filtered Mixture-of-Generators for Survival analysis), a two-level pipeline that reframes synthetic-data construction as sample selection rather than sample generation. A candidate pool is drawn from four architecturally distinct tabular generators, and each sample is scored by an ensemble of seven survival models trained on real data, using proper scoring rules as a per-sample plausibility proxy. An outer loop optimizes a selection policy—  \ngenerator quotas, scorer weights, a random complement, and stratified balancing on event time and censoring—against held-out downstream performance, while an inner loop tunes the downstream survival model (XGBoost-Cox). We evaluate FoGS on 16 public datasets under train-on-synthetic, test-on-real, reporting C-index and IBS on a 0–100 scale.  \nResults: FoGS yields mean improvements of +2 .17 in C-index and +0 .67 in IBS, improving both metrics on 9 of 16 datasets and at least one on 13 (one-sided Wilcoxon p = 0 .039 and p = 0 .035). It matches or exceeds real-data training on most cohorts, with no significant change in nearest-neighbour privacy margin relative to unfiltered sampling.  \nConclusion: Sample filtering over a heterogeneous generator pool is a viable substitute for real-data training in privacyrestricted clinical settings.  \nKeywords: Synthetic tabular data · Loss-guided filtering · Survival analysis · Concordance index · Generative models · Hyperparameter optimization · Integrated Brier score  \n1 Introduction  \nSurvival models guide clinical decision-making across oncology, cardiology, transplantation, and other specialties, supporting treatment stratification, follow-up scheduling, and clinical-trial design [15 , 24] . Their inputs are right-censored time-to-event data: tuples of covariates, observed time, and an event indicator, with the event time only partially known for subjects still under observation at study end [6 , 12] . Training survival models is constrained by structural properties of the data itself: clinically meaningful events accrue only over years of follow-up, which keeps cohort sizes small, and privacy regulations restrict sharing across institutions. These constraints make the curation of a sufficiently large training cohort the slowest and most expensive step in deploying a survival model in practice.  \nSynthetic tabular generation has been proposed as a remedy, with model fidelity steadily improving across the heterogeneous architectures developed for structured data. In principle, a high-quality synthetic cohort serves two complementary roles: augmenting scarce real-data training sets to improve downstream model robustness, and  \n∗ Corresponding author: [eugenio.lomurno@polimi.it](eugenio.lomurno@polimi.it).  \nenabling cohort sharing across institutions without disclosing patient-level records [27] . The viability of either role depends on whether synthetic data, when substituted for real training data, preserves the downstre","cbCaigHT6BlvY7Df","https://ap.wps.com/l/cbCaigHT6BlvY7Df","pdf",27492589,5,1,16,"English","en",105,"# Abstract\n# Introduction\n# Method Overview\n## Outer-loop Optimization\n## Inner-loop Survival Modeling\n# Results\n# Conclusion","[{\"question\":\"What problem does the paper address in survival model training?\",\"answer\":\"It targets the difficulty of training survival models with limited real cohorts, high costs of collecting time-to-event data over long follow-up, and privacy restrictions that prevent sharing across institutions.\"},{\"question\":\"How does FoGS build fully synthetic survival training data?\",\"answer\":\"FoGS draws candidate samples from a heterogeneous pool of four tabular generators, scores each sample using an ensemble of seven survival models trained on real data with proper scoring rules, and selects samples via an outer-loop policy.\"},{\"question\":\"What evidence shows FoGS works better than training on synthetic data from a single generator?\",\"answer\":\"Across 16 public datasets using train-on-synthetic and test-on-real, FoGS improves mean C-index and IBS, boosts both metrics on 9 datasets, and matches or exceeds real-data training on most cohorts without a significant change in nearest-neighbour privacy margin.\"}]",1784188073,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-filtered-mixture-of-generators-for-fully-synthetic-survival-training","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-filtered-mixture-of-generators-for-fully-synthetic-survival-training/83454/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in survival model training?","Question",{"text":76,"@type":77},"It targets the difficulty of training survival models with limited real cohorts, high costs of collecting time-to-event data over long follow-up, and privacy restrictions that prevent sharing across institutions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does FoGS build fully synthetic survival training data?",{"text":81,"@type":77},"FoGS draws candidate samples from a heterogeneous pool of four tabular generators, scores each sample using an ensemble of seven survival models trained on real data with proper scoring rules, and selects samples via an outer-loop policy.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence shows FoGS works better than training on synthetic data from a single generator?",{"text":85,"@type":77},"Across 16 public datasets using train-on-synthetic and test-on-real, FoGS improves mean C-index and IBS, boosts both metrics on 9 datasets, and matches or exceeds real-data training on most cohorts without a significant change in nearest-neighbour privacy margin.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]