[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84054-en":3,"doc-seo-84054-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84054,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Automating Quality Assessment of LLM-Generated Defeaters","High-integrity safety systems for autonomous vehicles and large-scale infrastructure rely on rooted directed acyclic graphs (DAGs) as structured assurance cases. To keep claims valid amid change, these cases must be robust against challenges called defeaters, whose LLM-generated quality is often judged manually. This work proposes an NLP method combining assurance-case graph structural analysis, vector semantic embeddings, and meta-classifiers trained on expert-consensus defeaters. Evaluations in automotive and energy domains show improved agreement (Cohen’s kappa) and an average F1-score of 0.84. The approach reduces subjective variance and supports scalable assurance-case automation.","Automating Quality Assessment of LLM-Generated  \nDefeaters  \nT. Rohlinger, D. Ratiu, and S. Wagner  \narXiv :2607 .06039v 1 [ cs . SE] 7 Jul 2026  \nAbstract—High-integrity systems, such as autonomous vehicle fleets or large-scale energy infrastructures, rely on structured assurance cases, represented as rooted directed acyclic graphs (DAGs), to justify safety claims. To remain valid in the face of change, these cases must be robust against potential challenges, known as ’defeaters’. While large language models (LLMs) have recently enabled scalable generation of such defeaters, validating their quality remains a predominantly manual and subjective process. This paper presents an automated method for assessing LLM-generated defeaters using natural language processing (NLP) techniques. Our approach combines structural analysis of assurance case graphs with vector-based semantic embeddings and meta-classifiers trained on expert-assessed consensus defeaters. We evaluate our method through two case studies in the automotive and energy domains, quantifying human reviewer dissensus using Cohen’s kappa (κ \u003C 0.442), which indicates low inter-rater agreement. Our automated approach achieves greater consistency with individual raters, improving (κ ≈ 40%). The method delivers an average F1-score of 0.84 across validation, reducing subjective variance through scalable, objective assessment. Our approach aims to advance the tool support for automation of assurance case synthesis.  \nIndex Terms—assurance case, defeater, NLP  \nI. INTRODUCTION  \nThe research community aims to establish a validation method for safety case fragments, assuring over time-changing safety-critical systems that adapt to evolving safety boundaries [1], [2] . The pace of adaptation required is determined by the exposure of entities in, for example, autonomous driving fleets. These kinds of systems demand robust safety cases, enhanced by runtime monitoring feedback [3], thereby enabling advanced system-level assurance [4] . Defeaters, in the domain of safety assurance, refer to arguments or evidence that undermine the assurance claims made within the safety case [5] . Identifying and understanding these defeaters are essential for validating a safety case’s resilience against realworld scenarios. Manual creation and validation in this context is not scalable and relies on subjective expert judgment [6] . Generating defeaters using large language models (LLMs) increases confidence in the argumentation [7], thereby making safety assurance scalable and responsive.  \nIn Viger et al. (2024) [7], the generated defeaters were manually reviewed by two safety experts. The experts only agreed on the quality of the defeaters, with a low level of  \n© 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. Published version: doi:10 . 1109/ICSRS68021 .2025.11422208.  \ninter-rater agreement calculated by cohens kappa [8] . Low inter-rater agreement indicates that expert judgments vary due to differing interpretations and domain knowledge. Furthermore, manual validation is resource-intensive and lacksscalability, particularly for dynamic systems with evolving safety boundaries, such as adaptive software. These limitations hinder the efficiency and reliability of assurance synthesis, making automated approaches necessary to enhance defeater validation. This work addresses the issue of subjectivity in manual defeater validation, as illustrated by the schematic figure 1 .  \nIn Section III, we propose an NLP-based method that uses BERT embeddings [9] and meta-classifiers to objectively evaluate LLM-generated defeaters, enabling scalable assurance. Our experiment, detailed","cbCaihPdQ51f7fcR","https://ap.wps.com/l/cbCaihPdQ51f7fcR","pdf",747469,4,1,10,"English","en",105,"# Introduction\n# Related Work\n# Method\n# Experiments and Results\n# Conclusion and Future Work","[{\"question\":\"What problem does the paper address about LLM-generated defeaters?\",\"answer\":\"It addresses that validating the quality of LLM-generated defeaters remains largely manual and subjective, which limits scalability and consistency for changing safety-critical systems.\"},{\"question\":\"How does the proposed automated method evaluate defeater quality?\",\"answer\":\"It combines structural analysis of assurance-case graphs, vector-based semantic embeddings, and meta-classifiers trained on expert-assessed consensus defeaters.\"},{\"question\":\"What evidence shows the method reduces subjectivity compared with human review?\",\"answer\":\"The evaluation across automotive and energy case studies quantifies low human inter-rater agreement (Cohen’s kappa) and shows the automated approach achieves greater consistency with individual raters, improving agreement by about 40%, with an average F1-score of 0.84.\"}]",1784192262,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"automating-quality-assessment-of-llm-generated-defeaters","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/automating-quality-assessment-of-llm-generated-defeaters/84054/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address about LLM-generated defeaters?","Question",{"text":75,"@type":76},"It addresses that validating the quality of LLM-generated defeaters remains largely manual and subjective, which limits scalability and consistency for changing safety-critical systems.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed automated method evaluate defeater quality?",{"text":80,"@type":76},"It combines structural analysis of assurance-case graphs, vector-based semantic embeddings, and meta-classifiers trained on expert-assessed consensus defeaters.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows the method reduces subjectivity compared with human review?",{"text":84,"@type":76},"The evaluation across automotive and energy case studies quantifies low human inter-rater agreement (Cohen’s kappa) and shows the automated approach achieves greater consistency with individual raters, improving agreement by about 40%, with an average F1-score of 0.84.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]