[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81677-en":3,"doc-seo-81677-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81677,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance","Medical large language models are positioned as near-term assistants rather than replacements, yet current evaluations typically measure isolated abilities such as clinical knowledge, EHR system interaction, or patient communication. Physician assistance demands coordinating these capabilities within a single, dynamic interaction where requests are underspecified, patient symptoms are ambiguous, and EHR tools require precise usage. PHYSASSISTBENCH is proposed as a benchmark for interactive doctor-patient-EHR assistance, built from real MIMIC-IV cases.","Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance  \nTianming Du1,2 , Peijie Yu3 , Sihan Shang1,2,4 , Danli Shi5 , My Linh Nguyen1,2 , Shengbo Gao1,2 , Guangyuan Li1,2 , Yinghong Yu1,2 , Yan Jiang1,6 , Qianlong Zhao2,7 , Behzad Bozorgtabar8 , Shaoxiong Ji1,9 , Jiazhen Pan10 , Daniel Rueckert10 , Jiancheng Yang1,2 * , 1ELLIS Institute Finland, 2Aalto University, 3Tencent, 4Harbin Institute of Technology, Shenzhen, 5Hong Kong Polytechnic University, 6University of Oulu, 7Polytechnic University of Milan, 8Aarhus University, 9University of Turku, 10Technical University of Munich  \n{du.tianming,[jiancheng.yang}@aalto.fi](jiancheng.yang}@aalto.fi)  \narXiv :2606 . 186 13v 3 [ cs .CL] 10 Jul 2026  \nAbstract  \nThe most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated capabilities: clinical knowledge, EHR system interaction, or patient communication. Physician assistance instead requires coordinating these capabilities within the same interaction, where physicians issue underspecified requests, patients describe symptoms ambiguously, and EHR systems demand precise tool use. We introduce PHYSASSISTBENCH, a benchmark for interactive doctor-patient-EHR assistance.  \nBuilt from real MIMIC-IV cases, PHYSASSISTBENCH uses a scalable pipeline to construct agentic patients: interactive, record-grounded agents that turn static EHR records into multiturn clinical scenarios while preserving clinical factuality. PHYSASSISTBENCH provides a curated bilingual evaluation set of 1,296 manually reviewed and physician-validated turns. Experiments with leading LLMs show that current models remain unreliable in this setting, which exposes a key bottleneck for clinical LLMs: reliable assistance requires coordination across knowledge, communication, and systems, not isolated gains in any of them.  \n1 Introduction  \nLLMs have generated substantial optimism for clinical AI (Thirunavukarasu et al., 2023 ; Moor et al., 2023 ; Rajpurkar et al., 2022) . Much of this optimism comes from their strong performance on medical examinations and question-answering benchmarks (Kung et al., 2023 ; Singhal et al., 2023), motivating visions of LLMs as a new “front door to healthcare”(NHS England-South East, 2024 ; Kyle, 2025) . However, recent studies suggest that such performance may not transfer to interactive clinical use. Laban et al. (2025) found that shifting from fully specified single-turn prompts to multi-turn, under-specified interactions caused an average 39%  \n* Corresponding author: Jiancheng Yang.  \nperformance drop. Similarly, Bean et al. (2026) found that LLMs performed strongly when tested alone, but failed to improve user performance in a randomized medical self-assessment study, identifying user interaction as a key barrier.  \nThese findings reveal a gap in current evaluation practice: practical failures may arise not only from insufficient medical knowledge, but from failures of interaction. As summarized in Table 1, prior medical LLM benchmarks mostly evaluate three isolated roles: knowledge, where models act as medical experts; system, where models retrieve or manipulate clinical records; and communication, where models interact with patients or generate clinical text. These roles are useful, but they miss the most plausible near-term deployment setting: assisting physicians under human oversight, as emphasized by ethical and regulatory expectations (World Health Organization, 2021 ; U.S. Food and Drug Administration, 2025) . Technically, this setting is also where recent studies expose a key bottleneck: interaction (Laban et al., 2025 ; Bean et al., 2026) . Physician assistance is not static question answering, but interactive coordination across incomplete physician intent, ambiguous patient information, and precise EHR actions.  \nFigure 1 illustrates the setting studied in this paper. Even for medical professionals, physician ","cbCaihZTOvzEUAwx","https://ap.wps.com/l/cbCaihZTOvzEUAwx","pdf",3112685,2,1,34,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does PHYSASSISTBENCH address in evaluating medical LLMs?\",\"answer\":\"It addresses evaluation gaps where existing benchmarks test isolated capabilities rather than the coordinated interaction needed for physician assistance across physician intent, patient dialogue, and EHR actions.\"},{\"question\":\"How is PHYSASSISTBENCH constructed?\",\"answer\":\"It is built from real MIMIC-IV cases using a scalable pipeline that creates agentic, record-grounded patients and multiturn clinical scenarios delivered via standardized FHIR interfaces.\"},{\"question\":\"What do experiments with current LLMs show on this benchmark?\",\"answer\":\"Current models remain unreliable in interactive doctor-patient-EHR assistance, highlighting coordination across knowledge, communication, and systems as a key bottleneck rather than isolated gains.\"}]",1784175357,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"are-llms-ready-to-assist-physicians-physassistbench-for-interactive-doctor-patient-ehr-assistance","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/are-llms-ready-to-assist-physicians-physassistbench-for-interactive-doctor-patient-ehr-assistance/81677/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does PHYSASSISTBENCH address in evaluating medical LLMs?","Question",{"text":75,"@type":76},"It addresses evaluation gaps where existing benchmarks test isolated capabilities rather than the coordinated interaction needed for physician assistance across physician intent, patient dialogue, and EHR actions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is PHYSASSISTBENCH constructed?",{"text":80,"@type":76},"It is built from real MIMIC-IV cases using a scalable pipeline that creates agentic, record-grounded patients and multiturn clinical scenarios delivered via standardized FHIR interfaces.",{"name":82,"@type":73,"acceptedAnswer":83},"What do experiments with current LLMs show on this benchmark?",{"text":84,"@type":76},"Current models remain unreliable in interactive doctor-patient-EHR assistance, highlighting coordination across knowledge, communication, and systems as a key bottleneck rather than isolated gains.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]