[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84880-en":3,"doc-seo-84880-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84880,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Collaborative Multi-Agent Testing for Emergent Failure Discovery in Autonomous Driving Systems","Autonomous Driving Systems (ADS) can fail due to faults inside individual modules and due to cross-module interactions across perception, planning, and control. Existing ADS testing often separates perturbation generation, behavioural assessment, and test exploration, reducing coordinated discovery of rare failures. CREAD introduces a collaborative multi-agent testing framework that coordinates these functions via a shared blackboard and orchestrator. In a work-in-progress setup, it uses perception fuzzing, metamorphic validation, and orchestration. Experiments in HighwayEnv show about 2.1× more failures per 100 scenarios than a single-agent baseline on average.","Collaborative Multi-Agent Testing for Emergent Failure Discovery in Autonomous Driving Systems  \nRuizhen Gu  \nQueen’s University Belfast Belfast, UK  \n[r.gu@qub.ac.uk](r.gu@qub.ac.uk)  \nKonstantinos Koufos  \nQueen’s University Belfast Belfast, UK  \n[k.koufos@qub.ac.uk](k.koufos@qub.ac.uk)  \nDonghwan Shin  \nThe University of Sheffield Sheffield, UK  \n[d.shin@sheffield.ac.uk](d.shin@sheffield.ac.uk)  \nVahid Garousi  \nQueen’s University Belfast Belfast, UK  \nAzerbaijan Technical University Azerbaijan [v.garousi@qub.ac.uk](v.garousi@qub.ac.uk)  \nMehrdad Dianati  \nQueen’s University Belfast Belfast, UK  \n[m.dianati@qub.ac.uk](m.dianati@qub.ac.uk)  \narXiv :2607 .06078v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Autonomous Driving Systems (ADS) can fail because of faults within individual modules as well as from interactions across perception, planning, and control. Yet existing ADS testing research often treats key testing functions, such as perturbation generation, behavioural assessment, and test case selection and exploration, as loosely coupled steps rather than coordinated roles for discovering such failures. We present CREAD, a collaborative multi-agent testing framework for testing ADS that organises perturbation generation, behavioural validation, and search coordination through a shared blackboard and an orchestrator. In the current work-in-progress instantiation, the framework focuses on perception-oriented perturbation generation, while remaining extensible to other ADS modules, including planning and control. It currently comprises a Perception Fuzzer Agent, a Metamorphic Validator Agent, and an Orchestrator Agent. Respectively, they generate perturbations, assess behavioural consistency across related scenario pairs, and coordinate further exploration. Experiments in HighwayEnv simulator show that the collaborative configuration improves failure discovery in the highway environment and remains competitive in the roundabout setting. Across the two environments, it yields about 2.1x as many failures per 100 scenarios as the single-agent baseline on average, while gains over a non-collaborative two-agent baseline vary across environments. These results suggest that collaborative multi-agent testing is a promising research direction for emergent ADS behaviour discovery.  \nIndex Terms—autonomous driving systems, software testing, multi-agent systems, large language models.  \nI. INTRODUCTION Autonomous Driving Systems (ADS) are safety-critical, open-world systems whose failures can arise both from faults within individual modules and from interactions among perception, planning, and control [1, 2] . Given this complexity, verification and validation for ADS increasingly relies on scenario-based testing to build safety-relevant evidence [3] .  \nIn this paradigm, functional scenarios derived from the operational design domain (ODD) are progressively refined into concrete scenarios, which can then be executed as test cases in simulation or other test environments [4] . Within such a  \npipeline, effective testing requires not only meaningful perturbation generation, but also reliable behavioural assessment and efficient exploration of the scenario space. Yet recent testing approaches based on feedback-guided fuzzing, adaptive search, and LLM-assisted scenario synthesis still struggle to consistently uncover rare but high-consequence failures, especially those arising in the long tail of safety-critical driving scenarios [5–7] . While some failures, such as collisions, are easy to recognise, more subtle behavioural degradations are harder to assess consistently, especially when they emerge from interactions across multiple modules [8] .  \nThis work is motivated by two limitations of current ADS testing. First, recent approaches of scenario generation, such as LLM-guided synthesis, improve realism and efficiency, but still struggle to steer exploration towards diverse and safetyrelevant scenario families without repeatedly generating simil","cbCaie1La22db7rX","https://ap.wps.com/l/cbCaie1La22db7rX","pdf",199397,1,6,"English","en",105,"# Introduction\n## Scenario-based testing for ADS\n## Limitations of existing pipelines\n## CREAD framework and collaborative roles\n## Closed-loop blackboard coordination","[{\"question\":\"What problem does CREAD target in autonomous driving system testing?\",\"answer\":\"CREAD targets the challenge of discovering rare, safety-relevant failures that arise from both module faults and interactions across perception, planning, and control, which existing pipelines often miss due to weak coupling between generation and validation.\"},{\"question\":\"How does CREAD coordinate multiple testing functions?\",\"answer\":\"CREAD organises perturbation generation, behavioural validation, and search coordination through a shared blackboard and an orchestrator, enabling continuous feedback between agents during exploration.\"},{\"question\":\"What agents are used in the current work-in-progress version?\",\"answer\":\"The current instantiation includes a Perception Fuzzer Agent for generating perception-oriented perturbations, a Metamorphic Validator Agent for comparing baseline and perturbed executions, and an Orchestrator Agent for prioritising scenario families for further testing.\"}]",1784198987,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"collaborative-multi-agent-testing-for-emergent-failure-discovery-in-autonomous-driving-systems","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/collaborative-multi-agent-testing-for-emergent-failure-discovery-in-autonomous-driving-systems/84880/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CREAD target in autonomous driving system testing?","Question",{"text":75,"@type":76},"CREAD targets the challenge of discovering rare, safety-relevant failures that arise from both module faults and interactions across perception, planning, and control, which existing pipelines often miss due to weak coupling between generation and validation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CREAD coordinate multiple testing functions?",{"text":80,"@type":76},"CREAD organises perturbation generation, behavioural validation, and search coordination through a shared blackboard and an orchestrator, enabling continuous feedback between agents during exploration.",{"name":82,"@type":73,"acceptedAnswer":83},"What agents are used in the current work-in-progress version?",{"text":84,"@type":76},"The current instantiation includes a Perception Fuzzer Agent for generating perception-oriented perturbations, a Metamorphic Validator Agent for comparing baseline and perturbed executions, and an Orchestrator Agent for prioritising scenario families for further testing.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]