[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81948-en":3,"doc-seo-81948-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81948,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Auto-DSM Under the Lens A Black-Box Evaluation Framework for LLM-Based DSM Generation","This paper proposes a black-box evaluation framework for assessing how well Large Language Models (LLMs) generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closed nature of existing Auto-DSM pipelines, the framework benchmarks generated DSMs against manually validated ground-truth matrices and supports single-run and multi-run evaluation. It combines structural, classification, and stability measures into a Composite Quality Score. Experiments on a fictive abstract system and a real refrigerator decomposition reveal sensitivity to ambiguity, inconsistent dependency definitions, and prompt formulation, clarifying hallucination and abstention failure modes while enabling auditable Auto-DSM benchmarking for MBSE integration.","Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation  \nN. Potters, T. Hofman  \narXiv :2607 .05985v 1 [ cs .AI ] 7 Jul 2026  \nAbstract—This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Design Structure Matrices (DSMs) from structured technical documentation. Motivated by the closedsource nature of current Auto-DSM pipelines, the framework introduces a reproducible methodology that benchmarks generated DSMs (GEN-DSMs) against manually validated groundtruth matrices (GT-DSMs). The evaluation integrates both singlerun and multi-run perspectives, combining structural metrics (Completeness, Correctness, Coupling Density), classification metrics (Selective Accuracy, Abstention Coverage), and stability measures (Entropy, Fleiss’ κ). To synthesize these aspects, a Composite Quality Score (Q) is proposed. Controlled experiments are conducted on two datasets: a fictive abstract system anda real-world refrigerator decomposition, covering variations in phrasing, parameterdataset alignment, and system complexity. Results show that LLMs can produce structurally plausible DSMsand achieve high reproducibility under well-structured inputs, but remain sensitive to ambiguity, inconsistent dependency definitions, and prompt formulation. The findings highlight systematic sources of hallucination and abstention failure, demonstrating both the potential and current limitations of LLM-driven DSM automation. The proposed framework provides a transparent benchmark for auditing Auto-DSM pipelines and establishes foundations for integrating LLM-based decomposition methods into model-based systems engineering (MBSE) workflows.  \nIndex Terms—Design Structure Matrix (DSM), System decomposition, Large Language Models (LLMs), Black-box evaluation, Model-Based Systems Engineering (MBSE), Dependency analysis, Reproducibility, Automation in engineering design  \nI. INTRODUCTION  \nThe growing complexity of engineered systems has driven demand for tools and methods that support system-level analysis and decision-making processes. Among these, the Design Structure Matrix (DSM), introduced by Steward [1], remains a foundational method for visualizing and analyzing interdependencies between system elements. Its compact N × N format supports modularity analysis, design optimization, and process structuring across domains such as product development and systems engineering [2] . Over decades, its use has matured through improvements in visualization and analysis toolse.g., hierarchical clustering and dependency algorithmsforming a robust framework for downstream decision support [2], [3] .  \nThe DSM-building phaseidentifying components and interactionsremains manual, resource-intensive, and error-prone  \nN. Potters and T. Hofman (e-mail: [t.hofman@tue.nl](t.hofman@tue.nl)) are with the Eindhoven University of Technology (TU/e), Dept. of Mechanical Engineering, Control Systems Technology section, Engineering Systems Design group, P.O.Box 513, 5600 MB Eindhoven, The Netherlands.  \n[2]–[5] . System decompositions typically rely on expert interviews and informal documentation, introducing variability, inconsistent terminology, and interpretation errors [2], [6], [7] . Poor documentation [8] and tacit organizational knowledge [7] further undermine reproducibility. As noted in [9], interviews are just as important as reading design documents, but they trade speed for accuracy, making DSM system decomposition costly and inconsistent.  \nModel-Based Systems Engineering (MBSE) provides a structured approach to managing system complexity and mitigating these issues [10] . By formalizing specifications in digital models, MBSE enables scalable and consistent dependency mapping [11]–[13] . A recent extension, the Elephant Specification Language (ESL) [14], extracts dependencies from structured natural language, improving traceability and reducing ambiguity. ESL support","cbCaijvXFTxXy0zZ","https://ap.wps.com/l/cbCaijvXFTxXy0zZ","pdf",2027709,4,1,29,"English","en",105,"# Introduction\n## Problem: Manual and error-prone DSM decomposition\n## MBSE and structured dependency extraction\n## LLMs for DSM generation and the need for evaluation","[{\"question\":\"What does the Auto-DSM black-box evaluation framework measure?\",\"answer\":\"It measures how accurately LLMs generate Design Structure Matrices from technical documentation by comparing generated DSMs to ground-truth matrices and evaluating structure, classification behavior, and stability.\"},{\"question\":\"How is evaluation performed in single-run and multi-run settings?\",\"answer\":\"The framework integrates both single-run and multi-run perspectives, using structural metrics, classification metrics such as selective accuracy and abstention coverage, and stability measures like entropy and Fleiss’ κ.\"},{\"question\":\"What factors most affect LLM performance in DSM generation?\",\"answer\":\"Results show sensitivity to ambiguity, inconsistent dependency definitions, and prompt formulation, which drive hallucination and abstention failures even when outputs look structurally plausible.\"}]","Auto-DSM Under the Lens A Black-Box Evaluation Framework for LLM-Based DSM Generation | PDF",1784177238,73,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"auto-dsm-under-the-lens-a-black-box-evaluation-framework-for-llm-based-dsm-generation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/auto-dsm-under-the-lens-a-black-box-evaluation-framework-for-llm-based-dsm-generation/81948/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does the Auto-DSM black-box evaluation framework measure?","Question",{"text":76,"@type":77},"It measures how accurately LLMs generate Design Structure Matrices from technical documentation by comparing generated DSMs to ground-truth matrices and evaluating structure, classification behavior, and stability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is evaluation performed in single-run and multi-run settings?",{"text":81,"@type":77},"The framework integrates both single-run and multi-run perspectives, using structural metrics, classification metrics such as selective accuracy and abstention coverage, and stability measures like entropy and Fleiss’ κ.",{"name":83,"@type":74,"acceptedAnswer":84},"What factors most affect LLM performance in DSM generation?",{"text":85,"@type":77},"Results show sensitivity to ambiguity, inconsistent dependency definitions, and prompt formulation, which drive hallucination and abstention failures even when outputs look structurally plausible.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]