[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133538-en":3,"doc-seo-133538-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},133538,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","CT-FineBench - A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation","Generated radiology report evaluation remains a critical barrier in CT report generation, because outputs contain long text, diverse and complex findings, and fine-grained, disease-oriented attributes. Existing metrics typically measure coarse lexical overlap or limited entity matching, masking clinically required diagnostic fidelity. CT-FineBench builds a QA-based benchmark grounded in CT-RATE and Merlin to score answer correctness for specific clinical attributes such as location, size, margin. Experiments show improved correlation with expert assessment and higher sensitivity to fine-grained factual errors than prior metrics.","CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation  \nRuifeng Yuan1,2,3 , Wanxing Chang1,2 , Weiwei Cao1,2,4 , Bowen Shi1,2,5 , Zhongyu Wei3 , Ling Zhang1 , Jianpeng Zhang1,2,4  \n1DAMO Academy, Alibaba Group, China, 2Hupan Lab, 310023, China,  \n3Fudan University, China, 4Zhejiang University, China, 5 Shanghai Jiao Tong University, China  \nCorrespondence: [jianpeng.zhang0@gmail.com](jianpeng.zhang0@gmail.com)  \nAbstract  \nThe evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of findings, and the presence of fine-grained, disease-oriented attributes. Conventional evaluation metrics offer only coarse measures of lexical overlap or entity matching and fail to reflect the granular diagnostic accuracy required for clinical use. To address this gap, we propose CT-FineBench, a benchmark built from CTRATE and Merlin to evaluate the fine-grained factual consistency of CT reports, constructed from CT-RATE and Merlin. Our benchmark is constructed through a meticulous, QuestionAnswering (QA) based process: first, we identify and structure key, finding-specific clinical attributes (e.g., location, size, margin) .  \nSecond, we systematically transform these attributes into a QA dataset, where questions probe for specific clinical details grounded in gold-standard reports. The evaluation protocol for CT-FineBench involves using this QA dataset to query a machine-generated report and scoring the correctness of the answers. This allows for a comprehensive, interpretable, and clinically-relevant assessment, moving beyond superficial lexical overlap to pinpoint specific clinical errors. Experiments show that CTFineBench correlates better with expert clinical assessment and is substantially more sensitive to fine-grained factual errors than prior metrics.  \n1 Introduction  \nThe automatic generation of radiology reports from medical images, particularly Computed Tomography (CT) scans, promises to enhance the efficiency of clinical workflows. However, the clinical adoption of such systems hinges on robust evaluation. For complex and information-dense CT reports, diagnostic fidelity is paramount. This extends beyond the identification of findings to the precise characterization of their clinical attributes. A single flaw  \nin reporting the location, morphology, or severity of a lesion can potentially compromise diagnostic accuracy. Therefore, a fine-grained evaluation metric focusing on clinical attributes is critical for CT report generation.  \nExisting evaluation metrics for radiology report generation can be classified into three types. Conventional linguistic evaluation metrics like ROUGE (Lin, 2004) and BLEU (Papineni et al., 2002), which are based on lexical overlap, are fundamentally inadequate for this task. Even more advanced, embedding-based metrics like BERTScore (Zhang et al., 2019), while better at capturing semantic similarity, still fail to identify and prioritize key medical information. Consequently, all these metrics often assign high scores to reports that are linguistically similar but clinically incorrect. Recognizing this gap, recent research has moved towards more clinically-aware evaluation paradigms. One type of work focuses on entity-based metrics, such as RadGraph (Jain et al., 2021) and RaTEScore (Zhao et al., 2024), which assess reports by extracting and comparing key medical entities like finding/disease and anatomical structures. However, their reliance on a limited set of coarse-grained entity and relation types means they often fail to capture the critical, fine-grained attributes that are crucial for diagnosis, particularly in complex CT reports. Another emerging approach, exemplified by metrics like GREEN (Ostmeier et al., 2024), leverages Large Language Models (LLMs) as judges. Their reliance on general LLMs turns them into black box that offers feedba","cbCaijw53X4BUVQv","https://ap.wps.com/l/cbCaijw53X4BUVQv","pdf",1326127,1,14,"English","en",105,"# Introduction\n## Why diagnostic fidelity is hard to evaluate\n## Limitations of existing metrics\n## QA-based evaluation motivation\n# CT-FineBench\n## Core idea and contributions\n## Benchmark construction workflow\n## Evaluation protocol (QA querying and scoring)","[{\"question\":\"Why are conventional metrics insufficient for CT report generation evaluation?\",\"answer\":\"They mostly rely on lexical overlap or semantic similarity and cannot reliably detect clinically critical mistakes in fine-grained diagnostic attributes.\"},{\"question\":\"How does CT-FineBench evaluate fine-grained factual consistency?\",\"answer\":\"It converts clinical finding attributes into a QA dataset, then queries a generated report and scores whether the answers to attribute-specific questions are correct.\"},{\"question\":\"What clinical attributes does CT-FineBench focus on?\",\"answer\":\"It emphasizes fine-grained attributes of findings, such as location, size, density, and margin, rather than only coarse findings.\"}]","CT-FineBench - A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation | PDF",1787221515,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"ct-finebench-a-diagnostic-fidelity-benchmark-for-fine-grained-evaluation-of-ct-report-generation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ct-finebench-a-diagnostic-fidelity-benchmark-for-fine-grained-evaluation-of-ct-report-generation/133538/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are conventional metrics insufficient for CT report generation evaluation?","Question",{"text":76,"@type":77},"They mostly rely on lexical overlap or semantic similarity and cannot reliably detect clinically critical mistakes in fine-grained diagnostic attributes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does CT-FineBench evaluate fine-grained factual consistency?",{"text":81,"@type":77},"It converts clinical finding attributes into a QA dataset, then queries a generated report and scores whether the answers to attribute-specific questions are correct.",{"name":83,"@type":74,"acceptedAnswer":84},"What clinical attributes does CT-FineBench focus on?",{"text":85,"@type":77},"It emphasizes fine-grained attributes of findings, such as location, size, density, and margin, rather than only coarse findings.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]