[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81536-en":3,"doc-seo-81536-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81536,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs","Evaluating retrieval-augmented generation (RAG) as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs). Three replicable EHR-based tasks are defined—imaging procedure extraction (modality, date, anatomic site), therapeutic antibiotic timeline generation, and key diagnosis identification—using real inpatient clinical notes from a US academic health system. Experiments compare multiple large language models with varying context sizes, contrasting targeted retrieval against recent-note input.","arXiv :2508 . 148 17v2 [ cs .CL] 9 Jul 2026  \nEvaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs  \nSkatje Myers 1 , Dmitriy Dligach2 , Timothy A. Miller3,4 , Samantha Barr 1 , James Landefeld 1 , Yanjun Gao5 , Matthew M. Churpek 1 , Anoop Mayampurath 1 , and Majid Afshar 1  \n1 University of Wisconsin-Madison  \n2 Loyola University Chicago  \n3 Boston Children’s Hospital  \n4 Harvard Medical School  \n5 University of Colorado-Anschutz  \nJuly 13, 2026  \nAbstract  \nObjective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records (EHRs) . Methods: We defined three EHR-based tasks that are replicable across health systems and vary in reasoning complexity: 1) extracting imaging procedures (modality, date, and anatomic site), 2) generating timelines of therapeutic antibiotic use, and 3) identifying the key diagnoses for a hospitalization. Using real inpatient clinical notes from a US academic health system, we evaluated three large language models (GPT-5.4-mini, Mistral Medium 3, DeepSeek V3.1) with varying amounts of provided context, comparing targeted retrieval to using the most recent clinical notes.  \nResults: For Imaging Procedures, RAG strongly outperformed recent-note inputs and exceeded longcontext performance (by 0.17-9.83 F1 across all models) using fewer than 8K tokens. Similar benefits were observed for Antibiotic Timelines, where \u003C8K of retrieved tokens matched long-context recent-notes performance (between-3.26 to +3.24 Jaccard) . Error analysis revealed that missing information in the clinical notes—often due to inter-hospital transfers—limited performance to some extent. However, performance on the Diagnosis Generation task remains largely static across methods and models.  \nDiscussion: RAG demonstrated strong token efficiency across tasks, with the clearest and most consistent gains observed for imaging extraction and antibiotic timeline reconstruction. Diagnosis generation proved the most challenging task, suggesting ceiling effects imposed by documentation variability and evaluation constraints.  \nConclusion: Our results suggest that RAG remains a competitive and efficient approach for clinical tasks over large amounts of EHR, even as newer models become capable of handling increasingly longer amounts of text.  \n1 Introduction  \nElectronic health records (EHRs) contain comprehensive documentation of patient care, including critical information for diagnosis and treatment planning. However, the volume of clinical notes has increased substantially in recent years, driven in part by copy-paste practices, templated documentation, and regulatory pressures—a phenomenon often referred to as “note bloat”. For example, nearly 1 in 5 patients arrive atthe emergency department with a chart the size of Moby Dick (more than 200K words) [1] . As a result, clinicians must navigate increasingly lengthy and redundant records to locate key information.  \nLarge language models (LLMs) can potentially alleviate this burden by assisting clinicians in quickly extracting information and reasoning over EHRs, and have demonstrated promising capabilities in clinical summarization [2] and question answering [3] . However, the sheer volume of clinical documentation can  \nexceed most LLMs’ context window size. A practical approach is to provide the most recent notes, which may suffice for some tasks but risks omitting crucial information buried in earlier documentation.  \nRetrieval-augmented generation (RAG) has emerged as a promising solution to using LLMs on long documents by retrieving only the most relevant text passages for a given task. Rather than processing entire patient charts, RAG systems can selectively extract pertinent clinical information to answer specific questions (Figure 1) . This approach can potentially reduce computational costs, improve accuracy by eliminating","cbCaihWb9QOtNQTz","https://ap.wps.com/l/cbCaihWb9QOtNQTz","pdf",1438208,1,20,"English","en",105,"# Abstract\n# Introduction\n## Note bloat and context limits\n## RAG motivation\n## Task definitions and replication\n# Methods\n# Results\n## Imaging Procedures\n## Antibiotic Timelines\n## Diagnosis Generation\n# Discussion\n# Conclusion","[{\"question\":\"What question does the study aim to answer about RAG versus long-context prompting?\",\"answer\":\"Whether retrieval-augmented generation can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records.\"},{\"question\":\"Which three EHR-based tasks are used to evaluate the approaches?\",\"answer\":\"Imaging procedure extraction, antibiotic timeline generation, and key diagnosis identification for a hospitalization.\"},{\"question\":\"How do RAG and long-context inputs compare in performance and token efficiency?\",\"answer\":\"RAG shows strong token efficiency with clear and consistent gains for imaging extraction and antibiotic timeline reconstruction, while diagnosis generation remains largely static and challenging across methods and models.\"}]",1784174122,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"evaluating-retrieval-augmented-generation-vs-long-context-input-for-clinical-reasoning-over-ehrs","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/evaluating-retrieval-augmented-generation-vs-long-context-input-for-clinical-reasoning-over-ehrs/81536/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the study aim to answer about RAG versus long-context prompting?","Question",{"text":75,"@type":76},"Whether retrieval-augmented generation can serve as an efficient alternative to long-context prompting for clinical reasoning over electronic health records.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which three EHR-based tasks are used to evaluate the approaches?",{"text":80,"@type":76},"Imaging procedure extraction, antibiotic timeline generation, and key diagnosis identification for a hospitalization.",{"name":82,"@type":73,"acceptedAnswer":83},"How do RAG and long-context inputs compare in performance and token efficiency?",{"text":84,"@type":76},"RAG shows strong token efficiency with clear and consistent gains for imaging extraction and antibiotic timeline reconstruction, while diagnosis generation remains largely static and challenging across methods and models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":28,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]