[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83890-en":3,"doc-seo-83890-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83890,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Toward Trustworthy Large Language Model Agents in Healthcare","Healthcare appointment scheduling remains a persistent operational bottleneck caused by manual coordination, fragmented legacy systems, and high administrative overhead, limiting provider availability and patient access. The paper introduces CareConnect, a safety-first conversational agent for healthcare logistics automation. It combines LLM function calling, retrieval-augmented generation (RAG), and layered deterministic safety guardrails, orchestrating eight domain tools for booking and facility information while refusing medical advice or diagnosis. Evaluation on 680 scenarios shows 91.8% task completion, 2.2s median latency, 96.0% safety compliance, and $0.0324 average cost per appointment.","Toward Trustworthy Large Language Model Agents  \nin Healthcare  \nHadi Hasan, Safaa Salman, Adam Tai Abou Dargham, Ammar Mohanna, Ali Chehab  \nDepartment of Electrical and Computer Engineering  \nAmerican University of Beirut  \nBeirut, Lebanon  \n[hsh24@mail.aub.edu](hsh24@mail.aub.edu), [sns44@mail.aub.edu](sns44@mail.aub.edu), [awt03@mail.aub.edu](awt03@mail.aub.edu), [am288@aub.edu.lb](am288@aub.edu.lb), [chehab@aub.edu.lb](chehab@aub.edu.lb)  \narXiv :2607 .05055v 1 [ cs .AI] 6 Jul 2026  \nAbstract—Healthcare appointment scheduling remains a persistent operational bottleneck, driven by manual coordination, fragmented legacy systems, and high administrative overhead. These inefficiencies constrain provider availability and degrade patient access to care. This paper presents CareConnect, a safetyfirst conversational agent for healthcare logistics automation that leverages large language model (LLM) function calling, retrieval-augmented generation (RAG), and layered deterministic safety guardrails. The system orchestrates eight domain-specific tools to support appointment booking, modification, cancellation, and facility information retrieval, while enforcing strict scope constraints that prohibit medical advice or diagnosis. Safetycritical situations are handled through deterministic short-circuit mechanisms for emergency detection and medical intent refusal. We evaluate CareConnect on a comprehensive benchmark of 680 task-oriented scenarios spanning end-to-end workflows, multiturn interactions, and edge cases. Experimental results demonstrate a 91.8% task completion rate with a median per-request latency of 2.2 seconds, 96.0% safety compliance on the dedicated safety-critical evaluation subset, and an average operational cost of $0.0324 per appointment, yielding a significant cost reduction compared to manual human scheduling. These findings show that carefully scoped and rigorously safeguarded LLM-based agents can reliably automate complex healthcare operational workflows while maintaining safety guarantees and achieving substantial cost efficiency. The source code and system implementation are publicly available at [https://github.com/Hadi-Hsn/CareConnect](https://github.com/Hadi-Hsn/CareConnect).  \nIndex Terms—large language models, agentic systems, healthcare logistics, conversational AI, retrieval-augmented generation, safety guardrails  \nI. INTRODUCTION  \nHealthcare appointment scheduling continues to impose a substantial administrative burden, with physicians spending an average of 16.6 hours per week on non-clinical tasks [1] . This overhead reduces clinical availability, increases operational costs, and limits timely patient access to care. At the same time, recent advances in large language models (LLMs) have demonstrated near expert-level performance on medical knowledge and reasoning benchmarks [2]–[4] . Despite these advances, most existing LLM systems are designed for clinical decision support or information retrieval and are not directly applicable to operational healthcare workflows such as appointment scheduling, which require strict safety, reliability, and transactional correctness [5] .  \nDeploying LLMs in healthcare logistics introduces several non-trivial challenges. First, safety and scope control are  \nparamount: conversational agents must not provide medical advice, perform diagnosis, or fail to appropriately handle emergency situations. Second, scheduling workflows depend on precise orchestration of stateful, transactional operations including availability search, booking, modification, and cancellation, where errors can result in patient harm and inconsistent system states. Third, operational interactions often require a combination of structured database actions and unstructured informational responses, whereas most RAG systems emphasize clinical knowledge retrieval rather than operational or logistical contexts [6], [7] . These constraints prevent the direct deployment of general-purpose LLMs in real","cbCaiaxjEr1Xase2","https://ap.wps.com/l/cbCaiaxjEr1Xase2","pdf",2472387,6,1,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the paper address in healthcare operations?\",\"answer\":\"It targets the administrative bottleneck of healthcare appointment scheduling, driven by manual coordination, fragmented legacy systems, and high overhead that reduces provider availability and delays patient access to care.\"},{\"question\":\"How does CareConnect enforce safety and scope during conversations?\",\"answer\":\"It uses deterministic pre-LLM intent filtering and scope-aware prompting, plus stateful tool-level validation. Emergency, diagnostic, and medical-advice requests are intercepted before LLM invocation, and safety-critical cases use deterministic short-circuit mechanisms with intent refusal.\"},{\"question\":\"What evaluation results does CareConnect achieve?\",\"answer\":\"On a benchmark of 680 task-oriented scenarios, CareConnect reaches a 91.8% task completion rate, 2.2 seconds median per-request latency, 96.0% safety compliance on safety-critical cases, and an average operational cost of $0.0324 per appointment.\"}]",1784191250,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"toward-trustworthy-large-language-model-agents-in-healthcare","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/toward-trustworthy-large-language-model-agents-in-healthcare/83890/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in healthcare operations?","Question",{"text":75,"@type":76},"It targets the administrative bottleneck of healthcare appointment scheduling, driven by manual coordination, fragmented legacy systems, and high overhead that reduces provider availability and delays patient access to care.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CareConnect enforce safety and scope during conversations?",{"text":80,"@type":76},"It uses deterministic pre-LLM intent filtering and scope-aware prompting, plus stateful tool-level validation. Emergency, diagnostic, and medical-advice requests are intercepted before LLM invocation, and safety-critical cases use deterministic short-circuit mechanisms with intent refusal.",{"name":82,"@type":73,"acceptedAnswer":83},"What evaluation results does CareConnect achieve?",{"text":84,"@type":76},"On a benchmark of 680 task-oriented scenarios, CareConnect reaches a 91.8% task completion rate, 2.2 seconds median per-request latency, 96.0% safety compliance on safety-critical cases, and an average operational cost of $0.0324 per appointment.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]