[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86092-en":3,"doc-seo-86092-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86092,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge","Medical vision-language pretraining (VLP) using paired CT images and radiology reports enables scalable clinical representation learning, yet existing approaches mostly align scans with whole reports or local regions with text fragments. OKA-CT leverages the anatomical organization of radiology findings by converting free-text reports into organ-conditioned knowledge via report parsing and LLM-assisted semantic structuring, then reuses the resulting organ hierarchy across two learning stages for contrastive and supervision signals. On CT-RATE and RAD-ChestCT, OKA-CT reaches zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, improving report-image alignment and sensitivity to disease-relevant regions.","Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report  \nKnowledge  \nGuoliang You 1 , Hongming Li 1 , Yuanwang Zhang2 , and Yong Fan 1,*  \n1 Department of Radiology, Perelman School of Medicine,  \nUniversity of Pennsylvania, Philadelphia, PA 19104, USA  \n2 School of Engineering and Applied Science,  \nUniversity of Pennsylvania, Philadelphia, PA 19104, USA  \narXiv :2607 . 10953v1 [ cs .CV] 12 Jul 2026  \nAbstract—Medical vision-language pretraining (VLP) from paired CT images and radiology reports enables scalable representation learning, but most existing methods align either whole scans with entire reports or local image regions with text fragments. These formulations underuse a key property of radiology reports: findings are organized around anatomical structures, with abnormalities described by organs, disease concepts, locations, and severity-related attributes. We propose OKA-CT, an organ-hierarchical knowledge-augmented framework for CT-report VLP. OKA-CT first converts free-text reports into organ-conditioned knowledge using radiology report parsing and LLM-assisted semantic structuring. The extracted hierarchy is used across two learning stages. Stage 1 injects anatomygrounded evidence into the CT visual representation through finegrained organ-conditioned supervision, while Stage 2 uses organspecific report evidence to guide structured report-CT contrastive learning, where hierarchy-derived semantic soft targets treat nonpaired cases with shared organ-level findings as weak semantic positives rather than uniform negatives. A lightweight querybased global branch further aggregates disease-relevant volumetric evidence for whole-scan representation. On CT-RATE and RAD-ChestCT datasets, OKA-CT achieves zero-shot abnormality diagnosis AUROCs of 84.9 and 72.2, outperforming prior CTVLP baselines. Retrieval and patch-occlusion analyses further show improved report-image alignment and stronger sensitivity to disease-associated anatomical regions.  \nIndex Terms—Medical Vision Language Pretraining, Organhierarchical Knowledge, Report Knowledge Extraction, Structured Contrastive Learning.  \nI. INTRODUCTION  \nMedical vision-language pretraining (VLP) provides a scalable paradigm for learning clinical visual representations from paired medical images and radiology reports. This paradigm is particularly valuable for computed tomography (CT), where volumetric studies are large, findings may be sparse, and dense expert annotations are expensive to obtain. Recent CT-report VLP methods have shown that routinely collected radiology reports can support zero-shot diagnosis, image-text retrieval, and general-purpose representation learning.  \nMost existing medical VLP methods organize image-report supervision at either the global or local level. Global alignment methods learn a single correspondence between an entire  \n*Corresponding author.  \nFig. 1. From coarse report alignment to organ-hierarchical evidence learning.(A) Existing CT VLP relies on global alignment or coarse local correspondence. (B) OKA-CT converts free-text reports into organ-hierarchical knowledge. (C) The hierarchy guides organ-conditioned supervision in Stage 1 and structured global and organ-level report-CT alignment in Stage 2 .  \nscan and its associated report, enabling large-scale contrastive learning from paired CT-report data and supporting zeroshot diagnosis and retrieval [1], [2] . Local and fine-grained methods further align image regions, report sentences, clinical entities, or anatomical structures, showing that spatially localized evidence is important for medical image understanding [3]–[8] . However, radiology reports are not simply wholescan descriptions or unordered collections of text fragments. They are naturally structured around anatomical systems: abnormalities are first grounded in organs, then described by disease or finding concepts, anatomical locations, severityrelated attributes, and diagnostic impressions","cbCaieemzJG5YbC9","https://ap.wps.com/l/cbCaieemzJG5YbC9","pdf",2805021,2,1,9,"English","en",105,"# Introduction\n## Global vs. local CT-report alignment\n## Limitations and remaining questions\n## Key observation and organ-conditioned knowledge\n## Proposed OKA-CT framework\n## Stage 1 and Stage 2 learning design","[{\"question\":\"What problem does OKA-CT address in CT vision-language pretraining?\",\"answer\":\"OKA-CT targets the limitation of existing global or local alignment methods that do not adequately use the organ-hierarchical structure of radiology reports, which can dilute sparse but clinically important findings.\"},{\"question\":\"How does OKA-CT transform radiology reports into organ-hierarchical knowledge?\",\"answer\":\"It parses free-text reports and uses LLM-assisted semantic structuring to normalize findings into organ-conditioned slots such as abnormality status, disease/concept entities, anatomical locations, and severity-oriented attributes.\"},{\"question\":\"What is the role of the organ hierarchy across OKA-CT’s two learning stages?\",\"answer\":\"Stage 1 injects anatomy-grounded evidence into CT visual representations with fine-grained organ-conditioned supervision, while Stage 2 guides structured report-CT contrastive learning using hierarchy-derived semantic soft targets, including weak positives for non-paired cases sharing organ-level findings.\"}]",1784208459,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-anatomy-grounded-ct-vision-language-representations-with-organ-hierarchical-report-knowledge","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learning-anatomy-grounded-ct-vision-language-representations-with-organ-hierarchical-report-knowledge/86092/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does OKA-CT address in CT vision-language pretraining?","Question",{"text":75,"@type":76},"OKA-CT targets the limitation of existing global or local alignment methods that do not adequately use the organ-hierarchical structure of radiology reports, which can dilute sparse but clinically important findings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does OKA-CT transform radiology reports into organ-hierarchical knowledge?",{"text":80,"@type":76},"It parses free-text reports and uses LLM-assisted semantic structuring to normalize findings into organ-conditioned slots such as abnormality status, disease/concept entities, anatomical locations, and severity-oriented attributes.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the organ hierarchy across OKA-CT’s two learning stages?",{"text":84,"@type":76},"Stage 1 injects anatomy-grounded evidence into CT visual representations with fine-grained organ-conditioned supervision, while Stage 2 guides structured report-CT contrastive learning using hierarchy-derived semantic soft targets, including weak positives for non-paired cases sharing organ-level findings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]