[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84159-en":3,"doc-seo-84159-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84159,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents","Conversational LLM agents can cause real-world harm when hidden workflow boundaries fail, such as missing identity checks or confirmation gates before taking irreversible actions. AGENTEVAL is a black-box testing framework that mines a conversational workflow graph from agent interactions, then targets specific stateful guards and prerequisites defined by the graph. It replays routes to each boundary, applies perturbations, and judges pass/fail using only dialogue turns. Benchmarks show 23–38 distinct boundary tests per agent versus a prompt-only baseline, with reduced duplicates and false alarms.","Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents  \nLiting Lin Boxi Yu  \nLero, the Research Ireland Centre for Software Lero, the Research Ireland Centre for Software  \nUniversity of Limerick, Ireland University of Limerick, Ireland  \n[Liting.Lin@ul.ie](Liting.Lin@ul.ie) [boxi.yu@lero.ie](boxi.yu@lero.ie)  \nYuzhong Zhang  \nThe Chinese University of Hong Kong, Shenzhen China  \n[123090848@link.cuhk.edu.cn](123090848@link.cuhk.edu.cn)  \nLionel Briand  \nLero, the Research Ireland Centre for Software University of Limerick, Ireland University of Ottawa, Canada [lionel.briand@lero.ie](lionel.briand@lero.ie)  \nDavid-Paul Niland Genesys Ireland  \n[davidpaulniland@gmail.com](davidpaulniland@gmail.com)  \nEmir Mu˜noz Genesys Ireland  \n[emir.munoz@gmail.com](emir.munoz@gmail.com)  \narXiv :2607 .06873v 1 [ cs . SE] 8 Jul 2026  \nAbstract—Conversational LLM agents can cause real-world harm when their internal workflows fail, such as completing a transaction without confirmation. Testing these state-dependent failures is difficult because critical boundaries, such as identity checks and confirmation gates, are hidden behind multiturn conversational prerequisites, rendering them inaccessible to standard tests. We present AGENTEVAL, a black-box testing framework that discovers and stresses these stateful boundaries. AGENTEVAL interacts with an agent to mine a conversational workflow graph, a model of its behavior. Instead of prompting blindly, AGENTEVAL uses this graph’s structure to enumerate specific guards and prerequisites as test targets, replaying the conversational path to a boundary before applying a perturbation. AGENTEVAL then executes each test, determining whether it passes or fails using only the conversation turns. We benchmark AGENTEVAL against a privileged, white-box auditor with access to the agent’s underlying source code, which AGENTEVAL never sees. On four τ3-bench agents, AGENTEVAL successfully generates tests covering 23–38 distinct boundaries per agent; ablation studies attribute the gain to the graph’s structure: 23 distinct boundaries versus 12 with a prompt-only baseline, at lower duplicate and false-alarm rates.  \nI. INTRODUCTION  \nLarge language models now power conversational agents that act on a user’s behalf [1], [2], calling tools to answer support requests, change bookings, and manage accounts [3] . However, conversational agents built on these models often fail, resulting in tangible real-world consequences [4],[5] . The most critical failures occur at the agent’s workflow boundaries: the points where a correct agent must enforce a guard, a prerequisite, or a validation before proceeding. A boundary that fails to hold changes the state of the system silently and often irreversibly: a booking canceled before the user confirmed it, an action taken for an unverified caller, a value accepted that should have been refused.  \nMost work on evaluating conversational LLM agents focuses on the underlying model’s capability, rather than the deployed system’s correctness: benchmarks such as τ-bench and  \nτ 2-bench [3], [6], WebArena [7], GAIA [8], AppWorld [9], and SWE-bench [10] score how well a model plans and completes tasks in a fixed harness. The deployed agent, however, is the model combined with the prompts, policies, tools, and guardrails wrapped around it, and a change to any of these can silently break a workflow on the next release. What the deployer needs is ordinary software testing of the deployed agent: generate tests, run them, and report faults repeatably. Testing these boundaries is hard. The agent is often a blackbox: reachable only through a conversational interface, with its internal prompts, tools, and state hidden [11] . If the testers do not have access to the internal state, they can see only the visible replies, and it is hard to tell where the boundaries are. These boundaries are also stateful, sitting behind multiturn prerequisites: a confirmation gate, for insta","cbCaicNPqLFOS40M","https://ap.wps.com/l/cbCaicNPqLFOS40M","pdf",389553,4,1,12,"English","en",105,"# Introduction\n## Problem: workflow boundary failures in conversational agents\n## Existing evaluation and why it falls short\n## AGENTEVAL approach: mining a conversational workflow graph\n## Graph-guided boundary testing and experiments","[{\"question\":\"What problem does AGENTEVAL address in conversational LLM agents?\",\"answer\":\"AGENTEVAL targets state-dependent failures at workflow boundaries, such as identity checks, confirmation gates, and other guards that are hidden behind multiturn conversational prerequisites.\"},{\"question\":\"How does AGENTEVAL find boundary locations without access to internal state?\",\"answer\":\"AGENTEVAL interacts with the agent to mine a conversational workflow graph, where nodes abstract agent activities and edges reflect observed user actions between them, making structural boundary locations discoverable.\"},{\"question\":\"How does AGENTEVAL run tests and decide whether they pass or fail?\",\"answer\":\"AGENTEVAL deterministically replays the conversational path to a chosen structural location, then applies a perturbation and determines pass/fail using only the conversation turns.\"}]",1784193522,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mining-workflow-graphs-for-black-box-boundary-testing-of-conversational-llm-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/mining-workflow-graphs-for-black-box-boundary-testing-of-conversational-llm-agents/84159/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does AGENTEVAL address in conversational LLM agents?","Question",{"text":75,"@type":76},"AGENTEVAL targets state-dependent failures at workflow boundaries, such as identity checks, confirmation gates, and other guards that are hidden behind multiturn conversational prerequisites.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does AGENTEVAL find boundary locations without access to internal state?",{"text":80,"@type":76},"AGENTEVAL interacts with the agent to mine a conversational workflow graph, where nodes abstract agent activities and edges reflect observed user actions between them, making structural boundary locations discoverable.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AGENTEVAL run tests and decide whether they pass or fail?",{"text":84,"@type":76},"AGENTEVAL deterministically replays the conversational path to a chosen structural location, then applies a perturbation and determines pass/fail using only the conversation turns.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]