[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82881-en":3,"doc-seo-82881-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82881,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Data-Driven Soft Labeling Scales for DNA Read Classification to Whole-Body Cell-Type Deconvolution","Cell-type deconvolution estimates the proportions of constituent cell types in heterogeneous biological samples, a central task in computational biology. DNA-methylation methods often rely on aggregated methylation measurements that discard per-read pattern information, and existing read-level classifiers do not scale due to non-discriminative reads dominating and hard labels conflicting with many-to-many read-to-cell-type mappings. The work proposes data-driven soft labels that estimate conditional cell-type distributions per read, integrating them into Syto, a modular read-level classification-based deconvolution framework. On a whole-body atlas of 39 human cell types, Syto reduces MSE by 2.56× versus state of the art and generalizes to an out-of-distribution dataset spanning 16 tissues.","arXiv :2607 .04987v2 [ cs .LG] 8 Jul 2026  \nData-Driven Soft Labeling Scales DNA Read Classification to Whole-Body Cell-Type Deconvolution  \nDmytro Rizdvanetskyi Nathan Roos Pavlo Lutsik  \nDepartment of Oncology  \nKU Leuven  \nLeuven, Belgium  \n{dmytro.rizdvanetskyi, nathanericjeanbaptiste.roos, [pavlo.lutsik}@kuleuven.be](pavlo.lutsik}@kuleuven.be)  \nAbstract  \nCell-type deconvolution, the task of estimating the proportions of constituent cell types in a heterogeneous biological sample, is a core problem in computational biology. Methods that rely on epigenetic marks such as DNA methylation typically operate on aggregated methylation estimates, discarding the pattern-level information carried by individual DNA reads. Existing read-level approaches that exploit this information are scarce, and all remain restricted to few-class settings; scaling them further is an open problem because, at scale, non-discriminative reads dominate and hard labels conflict with the many-to-many mapping between methylation patterns and cell types, preventing classifier convergence. To overcome this, we propose data-driven soft labels that estimate the conditional cell-type distribution for each read, and integrate this scheme into Syto, a new modular framework for read-level classification-based deconvolution. On a whole-body atlas of 39 human cell types, Syto reduces MSE by 2.56 × over SoTA, with gains transferring to an out-of-distribution dataset spanning 16 tissues. Syto lays the foundation for modeling increasingly large cell-type panels, with improved applications in biology and healthcare. The proposed soft-labeling scheme is further translatable to any setting with a many-to-many signal-to-label mapping.  \n1 Introduction  \nEpigenetic marks, such as histone modifications, chromatin accessibility and DNA methylation, collectively define cell identity and can be used to distinguish between different normal and malignant cells [1, 2] . Among these, DNA methylation stands out as the most stable and costeffective to profile at scale. Advances in methylome profiling have allowed researchers to exploit this biological signal for crucial clinical tasks: tumor classification, tumor purity estimation and cell-type deconvolution. Traditionally, these tasks are treated separately. Tumor classification scores a sample against predefined reference tumor cohorts to confirm a diagnosis and guide treatment [3, 4, 5] . In contrast, tumor purity and fraction estimation quantify the proportion of cancer cells in a tissue sample [6, 7] or circulating tumor DNA (ctDNA) in the bloodstream. While these estimates are vital for early  \ndiagnostics and monitoring disease progression [8], they often struggle to capture complex cellular heterogeneity, such as the specific makeup of the immune microenvironment [6] . Cell-type deconvolution addresses this gap by determining the precise proportions of all constituent cell types in  \nPreprint.  \nFigure 1: Principal scheme of the methylome-based cell-type deconvolution  \nan inhomogeneous sample. Because deconvolution fundamentally captures both the target tumor proportion and its specific cellular profile, the narrower problems of tumor classification and purity estimation can be advantageously reduced to it. Thus, accurate deconvolution covers a much broader range of clinical use cases than methods designed solely for classification or purity estimation. Existing methylome-based deconvolution methods typically operate on aggregated methylation measurements (β [9], α-values [10], absolute counts of methylated CpGs [11]) or fragment-level counts [12, 13, 14, 15], omitting the differences in methylation pattern at the level of the individual DNA fragments (or reads) that are known to be important for tumorigenesis [16] . Conversely, incorporating local omics signatures (potentially containing more than methylation) and the broader nucleotide context may offer a more accurate representation of cell-of-origin and may improve det","cbCaifUYn9l7CnIQ","https://ap.wps.com/l/cbCaifUYn9l7CnIQ","pdf",1461469,1,50,"English","en",105,"# Introduction\n# Related Work\n# Method Overview\n# Experiments and Results\n# Discussion","[{\"question\":\"Why do aggregated DNA methylation approaches limit read-level cell-type deconvolution?\",\"answer\":\"They typically use aggregated methylation estimates, discarding the pattern-level information embedded in individual DNA reads. This reduces the ability to capture discriminative within-read signals needed for accurate deconvolution.\"},{\"question\":\"What problem prevents scaling existing read-level classification methods to many cell types?\",\"answer\":\"At scale, non-discriminative reads dominate, and hard labels conflict with the many-to-many relationship between methylation patterns and cell types. This mismatch can stop classifier convergence and degrade performance.\"},{\"question\":\"How does the proposed Syto framework improve deconvolution accuracy across datasets?\",\"answer\":\"Syto introduces data-driven soft labels that estimate a conditional cell-type distribution for each read and integrates the scheme into a modular classification-based deconvolution framework. Experiments on a 39-cell-type whole-body atlas show a 2.56× MSE reduction over state of the art, with gains transferring to an out-of-distribution dataset covering 16 tissues.\"}]",1784183635,126,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"data-driven-soft-labeling-scales-for-dna-read-classification-to-whole-body-cell-type-deconvolution","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/data-driven-soft-labeling-scales-for-dna-read-classification-to-whole-body-cell-type-deconvolution/82881/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do aggregated DNA methylation approaches limit read-level cell-type deconvolution?","Question",{"text":75,"@type":76},"They typically use aggregated methylation estimates, discarding the pattern-level information embedded in individual DNA reads. This reduces the ability to capture discriminative within-read signals needed for accurate deconvolution.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem prevents scaling existing read-level classification methods to many cell types?",{"text":80,"@type":76},"At scale, non-discriminative reads dominate, and hard labels conflict with the many-to-many relationship between methylation patterns and cell types. This mismatch can stop classifier convergence and degrade performance.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed Syto framework improve deconvolution accuracy across datasets?",{"text":84,"@type":76},"Syto introduces data-driven soft labels that estimate a conditional cell-type distribution for each read and integrates the scheme into a modular classification-based deconvolution framework. Experiments on a 39-cell-type whole-body atlas show a 2.56× MSE reduction over state of the art, with gains transferring to an out-of-distribution dataset covering 16 tissues.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":21,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]