[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83122-en":3,"doc-seo-83122-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83122,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Reconfigurable Radiology Labels Without Relabeling","Public chest-radiograph (CXR) datasets often use fixed label schemas, yet free-text reports contain richer findings whose relevance varies by task, site, and reader. The work introduces a pipeline that converts radiology free-text reports into multi-label matrices via a single structured annotation pass, then reconfigures the label schema through dictionary edits rather than repeated corpus inference. Reconfiguring MIMIC-CXR cached annotations takes 196 seconds without API cost, versus thousands of dollars for full relabeling. A 58-label taxonomy shows 43% of studies include findings beyond CheXpert-14, improving label coverage.","arXiv :2607 .06597v1 [ ee ss .IV] 6 Jul 2026  \nReconfigurable Radiology Labels Without Relabeling  \nJean-Benoit Delbrouck Dave Van Veen  \nAkash Pattnaik Kalina Slavkova Javid Abderezaei  \nHarris Bergman Khan Siddiqui  \nHOPPR  \n§ [github.com/hopprai/radlabels](github.com/hopprai/radlabels)  \nPublic chest-radiograph (CXR) datasets are typically released with small, fixed label schemas such as CheXpert-14 . However, the underlying free-text reports describe far more findings—and which findings matter depends on the task, site, and reader. We release a pipeline that converts free-text reports into multi-label matrices with a single structured annotation pass, and then reconfigures the label schema through dictionary edits rather than new inference passes—i.e. , without re-parsing or relabeling the corpus. After this one-time pass, reconfiguring MIMIC-CXR (223K reports) from cached annotations takes 196 seconds with no API cost, compared to $6.6K for an equivalent relabeling pass with Claude Opus 4.7 . Using a 58-label taxonomy, we show that 43% of CXR studies contain at least one finding outside CheXpert-14 . Image probes trained on these labels match CheXpert-14 probes on shared targets while also reaching 0.78 AUROC on expert-reviewed long-tail labels that CheXpert-14 cannot represent. These results suggest a different unit of work for radiology labeling: once reports are structured, the label schema becomes a configuration to edit, not a corpus to relabel.  \n1. Introduction  \nRadiology image models are only as useful as the labels used to train and evaluate them. For many chestradiograph (CXR) classifiers, those labels are a multi-label vector: one entry per finding, marked resent, absent, or uncertain. Over the last decade, these vectors have usually been derived from radiology reports by automatic labelers. On CXR, the standard tools are rule-based systems such as NegBio (Peng et al. , 2018) and CheXpert-NLP (Irvin et al. , 2019), followed by BERT-based classifiers such as CheXbert (Smitet al. , 2020); all applied to the 14-finding CheXpert taxonomy. The same pattern appears beyond CXR: CT-RATE (Hamamci et al. , 2024) labels chest CT reports by annotating a small subset and distilling a RadBERT-style classifier to the remaining corpus. More recently, large language models (LLMs) have been used to extract broader label sets (Dorfner et al. , 2025), while entity-relation extractors such as RadGraph (Jain et al. , 2021 ; Delbrouck et al. , 2024) produce structured report annotations that downstream tools can consume.  \nThese developments have made large-scale radiology labeling practical, but most rely on fixed schemas. That is a lossy abstraction: reports describe devices, diseases, subtypes, uncertainty, anatomy, and site-specific phrasing that may not fit the schema. The right schema also depends on the task: one application may need pleural abnormality, another pleural effusion versus pleural thickening, and another laterality, severity, or temporal change. Beyond task dependencies, schema choices are also readerdependent: radiologists may disagree on the presence or description of a finding (Abujudeh et al. , 2010 ;  \nCorrespondence to: Jean-Benoit Delbrouck \u003C[jeanbenoit.delbrouck@hoppr.ai](jeanbenoit.delbrouck@hoppr.ai) >.  \nBruno et al. , 2015) . These sources of variation make flexibility a useful property of a labeling system; however, fixed taxonomies constrain what models learn and what researchers can measure.  \nThe practical barrier is schema iteration at corpus scale. Manual relabeling requires scarce radiologist time (McDonald et al. , 2015) . Distillation pipelines, such as the one used for CT-RATE, require anew annotation-and-training cycle when target labels change. LLMs can relabel reports more flexibly, but each schema change requires another inference pass over the corpus, introducing recurring cost, privacy constraints, and reproducibility concerns: one Claude Opus 4.7 pass over MIMIC-CXR costs approximatel","cbCaiiZK4Kghgo03","https://ap.wps.com/l/cbCaiiZK4Kghgo03","pdf",605898,1,23,"English","en",105,"# Introduction\n# Methods and Approach\n## Structured Report Annotator (SRA)\n## Radiological Aliases\n# Results and Evaluation\n# Discussion","[{\"question\":\"What evidence is provided for the effectiveness of the new labeling approach?\",\"answer\":\"Using a 58-label taxonomy, the paper reports that 43% of CXR studies contain at least one finding outside CheXpert-14. It also shows image probes trained on the reconfigured labels match CheXpert-14 on shared targets and achieve AUROC on long-tail labels not representable in CheXpert-14.\"}]",1784185425,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"reconfigurable-radiology-labels-without-relabeling","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/reconfigurable-radiology-labels-without-relabeling/83122/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What evidence is provided for the effectiveness of the new labeling approach?","Question",{"text":75,"@type":76},"Using a 58-label taxonomy, the paper reports that 43% of CXR studies contain at least one finding outside CheXpert-14. It also shows image probes trained on the reconfigured labels match CheXpert-14 on shared targets and achieve AUROC on long-tail labels not representable in CheXpert-14.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]