[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133537-en":3,"doc-seo-133537-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},133537,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","RadZero3D - Bridging Self-Supervised Video Models and Medical Vision-Language Alignment for Zero-Shot Chest CT Interpretation","The growing demand for automated analysis of 3D medical imaging, such as chest computed tomography (CT), highlights the need for generalizable and label-efficient vision-language models. However, extending vision-language models (VLMs) to volumetric data remains challenging due to limited annotated datasets and the complexity of aligning high-dimensional visual features with fine-grained clinical language. RadZero3D introduces voxel patch-text alignment for 3D chest CT using self-supervised video pretrained models, yielding strong zero-shot classification performance on internal and external validation settings.","This ICCV Workshop paper is the Open Access version, provided by the Computer Vision Foundation.  \nExcept for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.  \nRadZero3D: Bridging Self-Supervised Video Models and Medical Vision-Language Alignment for Zero-Shot Chest CT Interpretation  \nJonggwon Park* Kyoyun Choi* Byungmu Yoon Hong Geun Cho Bumcheol Hwang  \nDEEPNOID Inc.  \n{jgpark, kychoi, bmyoon, hgcho, [bchwang](bchwang}@deepnoid.com)[}](bchwang}@deepnoid.com)[@deepnoid.com](bchwang}@deepnoid.com)  \nAbstract  \nThe growing demand for automated analysis of 3D medical imaging, such as chest computed tomography (CT), highlights the need for generalizable and label-efficient visionlanguage models. However, extending vision-language models (VLMs) to volumetric data remains challenging due to the lack of large-scale annotated datasets and the complexity of aligning high-dimensional visual features with fine-grained clinical language. In this work, we introduce RadZero3D, a vision-language framework that brings fine-grained vision-language alignment to 3D chest CT utilizing self-supervised video pretrained models. Building on RadZero—a model originally designed for 2D chest X-rays—we adopt V-JEPA 2, a self-supervised video transformer, to encode volumetric inputs, and adapt its features to the medical domain using low-rank adaptation (LoRA). Text supervision is provided through LLMextracted “finding-sentences,” which describe localized observations in radiology reports. RadZero3D learns voxel patch-text associations using the VL-CABS mechanism of RadZero and a multi-positive contrastive loss. Trained on CT scans and radiology report pair dataset, RadZero3D achieves zero-shot classification performance comparable to existing 3D medical VLMs and demonstrates strong discriminative capability in both internal and external validation settings.  \n1. Introduction  \nMedical image analysis plays a critical role in clinical decision-making, enabling early diagnosis, treatment planning, and disease monitoring [16, 18] . Among various imaging modalities, 3D volumetric such as computed tomography (CT) and magnetic resonance imaging (MRI) provide significantly richer anatomical and pathological information compared to 2D images like chest X-rays [6, 23] .  \n*Equal contribution.  \nHowever, the increased complexity of 3D imaging necessitates detailed, labor-intensive analysis of volumetric scansand imposes a substantial burden on expert radiologists. Given the limited number of clinicians qualified to read 3Dscans, a substantial demand has emerged for automated assistance to alleviate the burden on radiologists and reduce the potential for human error [10, 12] .  \nIn line with the broader progress in AI, early approaches in medical image analysis largely relied on supervised learning [17, 25], which in turn depends on extensive labeled datasets. However, the high cost and limited availability of expert annotations in the medical domain have posed significant barriers to scalability. Recently, vision-language models (VLMs) have gained traction as a label-efficient alternative, enabling the use of image-text pairs instead of dense labels for training. Contrastive language-image pretraining (CLIP) [21] represents a foundational approach in VLM development, and its adaptation to the medical domain—leveraging medical images and their corresponding radiology reports—has yielded promising results [15, 28] .  \nWhile VLMs have been extensively applied to 2D modalities such as chest X-rays [5, 26, 27], their extension to 3D medical imaging remains relatively sparse. This can be attributed to the lack of large-scale public datasets and the limited availability of pretrained VLMs tailored for volumetric data. Existing approaches in this area are often limited to global image-text alignment [11] or rely on auxiliary segmentation models [24], which restrict their generalizability","cbCaic3MAuHtepLA","https://ap.wps.com/l/cbCaic3MAuHtepLA","pdf",7907067,3,1,"English","en",105,"# Abstract\n# Introduction\n# Related Works\n## Vision-Language Models for 3D Medical Imaging","[{\"question\":\"What problem does RadZero3D address in 3D medical vision-language modeling?\",\"answer\":\"RadZero3D targets the difficulty of extending vision-language models from 2D to volumetric 3D CT, where large-scale labeled datasets are scarce and fine-grained text alignment is complex.\"},{\"question\":\"How does RadZero3D create alignment between CT images and radiology text?\",\"answer\":\"It uses self-supervised video pretrained transformer features (V-JEPA 2) adapted to medical data with LoRA, then learns voxel patch-text associations via RadZero’s VL-CABS mechanism and a multi-positive contrastive loss.\"},{\"question\":\"What datasets and supervision are used to train and evaluate RadZero3D?\",\"answer\":\"Training relies on CT scans paired with radiology reports, using LLM-extracted finding sentences for text supervision, and evaluation is performed with internal and external validation on public datasets including CT-RATE.\"}]","RadZero3D - Bridging Self-Supervised Video Models and Medical Vision-Language Alignment for Zero-Shot Chest CT Interpretation | PDF",1787221501,20,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"radzero3d-bridging-self-supervised-video-models-and-medical-vision-language-alignment-for-zero-shot-chest-ct-interpretation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/radzero3d-bridging-self-supervised-video-models-and-medical-vision-language-alignment-for-zero-shot-chest-ct-interpretation/133537/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-27","2026-08-20",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RadZero3D address in 3D medical vision-language modeling?","Question",{"text":75,"@type":76},"RadZero3D targets the difficulty of extending vision-language models from 2D to volumetric 3D CT, where large-scale labeled datasets are scarce and fine-grained text alignment is complex.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RadZero3D create alignment between CT images and radiology text?",{"text":80,"@type":76},"It uses self-supervised video pretrained transformer features (V-JEPA 2) adapted to medical data with LoRA, then learns voxel patch-text associations via RadZero’s VL-CABS mechanism and a multi-positive contrastive loss.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets and supervision are used to train and evaluate RadZero3D?",{"text":84,"@type":76},"Training relies on CT scans paired with radiology reports, using LLM-extracted finding sentences for text supervision, and evaluation is performed with internal and external validation on public datasets including CT-RATE.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":29,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":29,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]