[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86450-en":3,"doc-seo-86450-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86450,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences","Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, yet hallucinations create serious safety and decision risks. The study presents a benchmark framework to measure the faithfulness of LLM-generated trial summaries for three stakeholder audiences using 200 stratified trials, audience-specific prompts, and a six-dimension annotation scheme. Baselines are reported for GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash, and a knowledge-graph-augmented retrieval system grounded in PubMed improves NLI-based faithfulness with statistically significant gains.","Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences  \nRobert Williams  \nUniversity of Texas at Austin  \nComputer & Data Science Online  \nAustin, Texas, USA  \n[rgw@utexas.edu](rgw@utexas.edu)  \narXiv :2607 .09932v 1 [ cs .CL] 10 Jul 2026  \nAbstract  \nLarge language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their tendency to hallucinate poses significant risks in this high-stakes context. This study introduces a benchmark evaluation framework for measuring the faithfulness of LLM-generated clinical trial summaries across three stakeholder audiences. The framework consists of 200 stratified trials drawn from the Aggregate Analysis of ClinicalTrials.gov database, evaluated using audiencespecific prompt templates and a six-dimension faithfulness annotation schema. Baseline measurements were established for GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash across 1,800 generated summaries scored using a cross-encoder natural language inference (NLI) model. Unsupported Claims was identified as the dominant failure mode across all three models, with a mean annotation score of 1.55 out of three. A knowledge-graph-augmented retrieval system grounded in the PubMed Knowledge Graph was developed and evaluated against the baseline, producing statistically significant improvements in NLI-based faithfulness scores (entailment +0 . 0125, faithfulness +0 .0130, 􀀿 \u003C 0. 0001) . Improvement pathways were model-dependent, with GPT-4o improving primarily through contradiction reduction while Claude Sonnet 4.6 and Gemini 2.5 Flash improved through increased entailment.  \nCCS Concepts  \n• Computing methodologies → Natural language processing;  \n• Applied computing → Health informatics.  \nKeywords  \nclinical trial summarization, faithfulness evaluation, hallucination detection, knowledge graph, retrieval-augmented generation, NLI, large language models  \nACM Reference Format:  \nRobert Williams. 2026. Faithful by Design: Evaluating and Improving LLMGenerated Clinical Trial Summaries for Multi-Stakeholder Audiences. In Proceedings of UT Austin AI in Health (AI in Health ’26) . ACM, New York, NY, USA, 8 pages. [https://doi.org/10.1145/nnnnnnn.nnnnnnn](https://doi.org/10.1145/nnnnnnn.nnnnnnn)  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission [and/or a fee. Request permissions from permissions@acm.org](and/or a fee. Request permissions from permissions@acm.org).  \nAI in Health ’26, Austin, TX  \n© 2026 Copyright held by the owner/author(s) . Publication rights licensed to ACM. ACM ISBN 978-x-xxxx-xxxx-x/YYYY/MM [https://doi.org/10.1145/nnnnnnn.nnnnnnn](https://doi.org/10.1145/nnnnnnn.nnnnnnn)  \n1 Introduction  \nClinical trial results are among the most consequential documents in healthcare. They underpin regulatory approvals, inform prescribing and treatment decisions, guide coverage determinations by payers, and must be communicated to patients in plain language to support informed consent. A single completed trial may require distinct summaries for healthcare providers, patients, and payers, each demanding a different framing of the same underlying evidence. ClinicalTrials.gov currently lists more than 400,000 registered trials, and this summarization workload currently falls to skilled medical writers who must translate complex findings for each audience without distorting them.  \nLarge language models (LLMs) are natural candidates for automating this ","cbCaiaki1mSgBiSs","https://ap.wps.com/l/cbCaiaki1mSgBiSs","pdf",1930797,2,1,"English","en",105,"# Introduction\n## Motivation and stakes in clinical communication\n## Related work and limitations of existing metrics\n## Contributions and overview of the proposed framework","[{\"question\":\"Why is faithfulness critical when using LLMs for clinical trial summaries?\",\"answer\":\"Clinical trial summaries influence regulatory approvals, prescribing decisions, payer coverage determinations, and patient understanding. Hallucinated or distorted claims can mislead stakeholders and lead to harmful outcomes.\"},{\"question\":\"What does the proposed benchmark evaluation framework include?\",\"answer\":\"The framework uses 200 stratified trials from the AACT database, evaluates summaries for three stakeholder audiences with audience-specific prompt templates, and applies a six-dimension faithfulness annotation schema on a one-to-three ordinal scale.\"},{\"question\":\"Which failure mode dominated across the evaluated LLMs, and how was improvement achieved?\",\"answer\":\"Unsupported Claims was the dominant failure mode across GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash based on annotation scores. A knowledge-graph-augmented retrieval system grounded in PubMed Knowledge Graph improved NLI-based faithfulness scores with statistically significant gains, with improvement pathways differing by model.\"}]",1784211811,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"faithful-by-design-evaluating-and-improving-llm-generated-clinical-trial-summaries-for-multi-stakeholder-audiences","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/faithful-by-design-evaluating-and-improving-llm-generated-clinical-trial-summaries-for-multi-stakeholder-audiences/86450/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is faithfulness critical when using LLMs for clinical trial summaries?","Question",{"text":74,"@type":75},"Clinical trial summaries influence regulatory approvals, prescribing decisions, payer coverage determinations, and patient understanding. Hallucinated or distorted claims can mislead stakeholders and lead to harmful outcomes.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What does the proposed benchmark evaluation framework include?",{"text":79,"@type":75},"The framework uses 200 stratified trials from the AACT database, evaluates summaries for three stakeholder audiences with audience-specific prompt templates, and applies a six-dimension faithfulness annotation schema on a one-to-three ordinal scale.",{"name":81,"@type":72,"acceptedAnswer":82},"Which failure mode dominated across the evaluated LLMs, and how was improvement achieved?",{"text":83,"@type":75},"Unsupported Claims was the dominant failure mode across GPT-4o, Claude Sonnet 4.6, and Gemini 2.5 Flash based on annotation scores. A knowledge-graph-augmented retrieval system grounded in PubMed Knowledge Graph improved NLI-based faithfulness scores with statistically significant gains, with improvement pathways differing by model.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]