[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86427-en":3,"doc-seo-86427-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86427,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Learning from Lost Provenance Multiple Instance Learning for Cancer Registry Tumor Group Classification","Modernizing cancer registries with deep learning enables automation of labor-intensive tasks such as coding pathology reports, yet progress is limited by the lack of report-level human annotations. Cancer registries produce many operational expert labels at the patient level, but the mapping to the specific pathology reports that generated them is not retained. This work trains efficient deep learning classifiers without per-report labeling by recovering the lost linkage using ABMIL for tumor group classification at the BC Cancer Registry.","arXiv :2607 .0348 1v2 [ cs .CL] 11 Jul 2026  \nLearning from Lost Provenance: Multiple Instance Learning for Cancer Registry Tumor Group Classification  \nLeonard Ruocco 1 , Jonathan Simkin 1,2 , Lovedeep Gondara2 , Gregory Arbour3 , and  \nRaymond Ng3  \n1 British Columbia Cancer Registry, Provincial Health Services Authority, Vancouver, Canada  \n2 University of British Columbia, Vancouver, Canada  \n3 Data Science Institute, University of British Columbia, Vancouver, Canada  \nJuly 14, 2026  \nAbstract  \nModernizing cancer registries with deep learning is opening new opportunities to automate labor-intensive tasks such as the coding of pathology reports. However, progress is constrained by the scarcity of report-level human-annotated training data. Cancer registries generate substantial volumes of expert-assigned labels as a routine product of their operations, but these exist at the patient level and are not linked to the individual pathology reports that informed them, limiting their direct use for training models. We develop an efficient framework for training deep learning classifiers by leveraging these operationally-generated labels without requiring per-report human annotation, demonstrated for tumor group classification at the BC Cancer Registry. We use Attention-Based Multiple Instance Learning (ABMIL) to recover the lost link between patient-level labels and the reports that informed them, leveraging the attention the model places on each report to distil a large, noisily-labeled corpus into a compact, high-quality per-report training dataset. A classifier fine-tuned on a distilled dataset achieved a macro F1 of 0.83, outperforming established baselines across most tumor groups. By turning routine operational labels into high-quality training data without additional annotation or large-scale computing infrastructure, ABMIL offers a practical and accessible route to automating cancer registry workflows.  \n1 Introduction  \nPopulation-based cancer registries (PBCRs) play a foundational role in cancer surveillance, informing public health policy, resource allocation, and epidemiological research [1, 2] . A core operational task in any PBCR is the assignment of tumor group — a coded classification of cancer type used to direct cases to the appropriate tumor-specific coding workflows and populate surveillance databases [1] . At the BC Cancer Registry (BCCR), one of Canada’s largest provincial cancer registries, this task involves classifying incoming pathology reports into oneof many standardized tumor groups spanning the full spectrum of malignancy. The volume of incoming reports is substantial: BCCR receives approximately 100,000 reportable pathology reports annually, and the manual effort required to assign tumor group classifications represents a significant and growing burden [3, 4] on certified Oncology Data Specialists (ODS)—trained health information management professionals who code cancer cases into registry databases.  \nNatural language processing (NLP) offers a principled route to automating classification of pathology reports [5] . Pathology reports are semi-structured free-text documents with a consistent clinical vocabulary and relatively constrained linguistic variation, making them wellsuited to supervised text classification [6] . Recently, deep learning approaches have supplanted traditional rule-based and statistical machine learning (ML) methods for clinical NLP tasks [7], with transformer-based models now representing the state of the art [8] . Central to these approaches is the representation of text as dense embedding vectors that capture the semantic content and contextual relationships between words [9] .  \nA central challenge in developing NLP systems for cancer registry automation is the scarcity of pathology reports labeled at the individual document level in a format suitable for supervised ML. Constructing such datasets requires substantial subject matter expert involvement; as an ODS must review ","cbCaiomYtSazIMiq","https://ap.wps.com/l/cbCaiomYtSazIMiq","pdf",1024322,1,11,"English","en",105,"# Introduction\n## Challenge: lack of report-level annotations\n## Operational labels and the missing report-to-label link\n## Existing strategies and their limitations\n## Proposed approach: ABMIL framework","[{\"question\":\"What problem does the paper address in cancer registry automation?\",\"answer\":\"It addresses the lack of report-level human-labeled pathology data needed for supervised NLP models in cancer registries, while operational labels exist only at the patient level without a recorded link to the underlying reports.\"},{\"question\":\"How does the method use patient-level labels without per-report annotation?\",\"answer\":\"It recovers the missing linkage between patient labels and the reports that informed them by using Attention-Based Multiple Instance Learning (ABMIL), which leverages attention over reports to distill a large noisily labeled corpus into compact per-report training data.\"},{\"question\":\"How is performance evaluated and what result is reported?\",\"answer\":\"A classifier fine-tuned on the distilled dataset achieves a macro F1 of 0.83 and outperforms established baselines across most tumor groups, demonstrating improved effectiveness for tumor group classification.\"}]",1784211686,28,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"learning-from-lost-provenance-multiple-instance-learning-for-cancer-registry-tumor-group-classification","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/learning-from-lost-provenance-multiple-instance-learning-for-cancer-registry-tumor-group-classification/86427/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in cancer registry automation?","Question",{"text":75,"@type":76},"It addresses the lack of report-level human-labeled pathology data needed for supervised NLP models in cancer registries, while operational labels exist only at the patient level without a recorded link to the underlying reports.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method use patient-level labels without per-report annotation?",{"text":80,"@type":76},"It recovers the missing linkage between patient labels and the reports that informed them by using Attention-Based Multiple Instance Learning (ABMIL), which leverages attention over reports to distill a large noisily labeled corpus into compact per-report training data.",{"name":82,"@type":73,"acceptedAnswer":83},"How is performance evaluated and what result is reported?",{"text":84,"@type":76},"A classifier fine-tuned on the distilled dataset achieves a macro F1 of 0.83 and outperforms established baselines across most tumor groups, demonstrating improved effectiveness for tumor group classification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]