[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84000-en":3,"doc-seo-84000-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84000,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Articulating Assumptions in AI-Generated Scientific Analyses through Task Decomposition","Scientific results produced by LLM-generated analysis code must be understandable and reproducible, yet uncertainty can emerge both from the original natural-language specification and from the stochastic code generation process. Even executable code may fail to clarify which quantities are computed and which assumptions drive final outcomes. A multi-agent framework is introduced to ground quantities via semantic differencing, separating code generation, execution, tracing, and validation across agents. An ambiguity-inspection module proposes instruction rewrites before generation. Validation on collider physics analyses shows improved transparency and reliability versus single-prompt approaches and supports using substantially smaller models.","arXiv :2607 .05762v 1 [ cs . SE] 7 Jul 2026  \nArticulating Assumptions in AI-Generated Scientific Analyses  \nthrough Task Decomposition  \nAhmed Hammada∗ and Mihoko Nojirib,c,d†  \na Center of AI and Natural science, KIAS, Seoul 02455, Korea.  \nb Theory Center, IPNS, KEK, 1-1 Oho, Tsukuba, Ibaraki 305-0801, Japan.  \nc The Graduate University of Advanced Studies (Sokendai), 1-1 Oho, Tsukuba, Japan. d Kavli IPMU (WPI), University of Tokyo, 5-1-5 Kashiwanoha, Kashiwa, Chiba 277-8583,  \nJapan.  \nAbstract  \nScientific results produced by LLM generated analysis code must be understandable and reproducible. However, uncertainty can arise at different stages of the process, both in the original natural language specification and in the generated implementation. As a result, even executable code may not provide a clear understanding of which quantities are being computed or which assumptions determine the final results. To address this challenge, we introduce quantity grounded semantic differencing, a multi-agent framework for analyzing and comparing scientific programs generated by LLMs. The framework assigns code generation, execution, tracing, and validation to separate agents, allowing it to reconstruct how key output quantities are produced and to identify differences between the intended analysis and the implemented code. We also introduce a module that inspects ambiguities in the initial user instruction and suggests alternative rewrites before code generation. Its modular design enables application to different scientific domains by replacing domain specific resources while preserving the same workflow. We validate the framework on representative collider physics analyses. The results demonstrate that the modular task decomposition enhances both transparency and reliability relative to the previous single prompt approach, while enabling substantially smaller models to execute the complete workflow.  \n∗ email: [hammad@kias.re.kr](hammad@kias.re.kr)  \n†email: [nojiri@post.kek.jp](nojiri@post.kek.jp)  \nPackage: GitHub  \nContents  \n1 Introduction 2  \n2 Multi-Agent Framework 3  \n2.1 Methodology 4  \n3 A Case Study of Domain Specific Tuning 7  \n3.1 Code generation profile 7  \n3.2 Helper registry and helper selection policy 8  \n3.3 Critique context 8  \n3.4 Oracle context 8  \n4 Workflow Validation 9  \n4.1 Preparation of benchmark task cards 9  \n4.2 Stability of helper selection 10  \n4.3 Validation of generated code 11  \n4.4 Ambiguity Detection 12  \n4.5 Tracer and Critique 14  \n5 Conclusion 17  \n1 Introduction  \nLarge language models (LLMs) have recently achieved remarkable progress in generating executable code from natural language descriptions [1–3] . This capability is beginning to transform scientific computing by assisting not only in the development of new software but also in the modification of existing code and the construction of complex analysis pipelines [4] . As a result, researchers no longer need to implement every computational step manually. Instead, they can describe computational tasks in natural language and integrate the generated code into a broader scientific workflow [5] .  \nHowever, natural language based code generation introduces two distinct sources of uncertainty. The first arises from the request itself. Unlike programming languages, natural language does not impose a strict specification of computational procedures and often leaves room for interpretation in essential aspects of an analysis, including the target set of objects, exclusion criteria, aggregation units, and derivation rules [6] . The second arises from the variability and fallibility of the stochastic generation process. Even for identical or closely related requests, a language model may produce different outputs, some of which may be inappropriate, incomplete, or semantically incorrect [7] . Consequently, the executability of generated code does not ensure that users understand what the output represents, nor does it make explicit the assumpti","cbCaiby3akbc7UMC","https://ap.wps.com/l/cbCaiby3akbc7UMC","pdf",543040,5,1,20,"English","en",105,"# Introduction\n# Multi-Agent Framework\n## Methodology\n# A Case Study of Domain Specific Tuning\n## Code generation profile\n## Helper registry and helper selection policy\n## Critique context\n## Oracle context\n# Workflow Validation\n## Preparation of benchmark task cards\n## Stability of helper selection\n## Validation of generated code\n## Ambiguity Detection\n## Tracer and Critique\n# Conclusion","[{\"question\":\"Why can LLM-generated scientific analysis code be insufficient for reproducibility?\",\"answer\":\"Uncertainty can originate in the natural-language specification and in the stochastic nature of code generation. As a result, executable code may not clearly reveal which quantities are computed or what assumptions determine final results.\"},{\"question\":\"What is the purpose of quantity grounded semantic differencing in this framework?\",\"answer\":\"It reconstructs how key output quantities are produced by assigning code generation, execution, tracing, and validation to separate agents. This enables differences between the intended analysis and the implemented code to be identified.\"},{\"question\":\"How does the framework handle ambiguities in the initial user instruction?\",\"answer\":\"A dedicated module inspects ambiguities in the user request and suggests alternative rewrites before code generation, improving alignment between the specification and the implemented analysis.\"}]",1784191951,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"articulating-assumptions-in-ai-generated-scientific-analyses-through-task-decomposition","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/articulating-assumptions-in-ai-generated-scientific-analyses-through-task-decomposition/84000/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why can LLM-generated scientific analysis code be insufficient for reproducibility?","Question",{"text":76,"@type":77},"Uncertainty can originate in the natural-language specification and in the stochastic nature of code generation. As a result, executable code may not clearly reveal which quantities are computed or what assumptions determine final results.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the purpose of quantity grounded semantic differencing in this framework?",{"text":81,"@type":77},"It reconstructs how key output quantities are produced by assigning code generation, execution, tracing, and validation to separate agents. This enables differences between the intended analysis and the implemented code to be identified.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the framework handle ambiguities in the initial user instruction?",{"text":85,"@type":77},"A dedicated module inspects ambiguities in the user request and suggests alternative rewrites before code generation, improving alignment between the specification and the implemented analysis.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]