[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83664-en":3,"doc-seo-83664-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83664,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","AutoResearch An Execution Grounded Multi Agent Framework for Reliable Research Workflow Automation","Automated research agents generate code, retrieve literature, and draft scientific artifacts, yet often lack a validation loop covering whether experiments execute correctly and whether sources support stated claims. AutoResearch presents an execution-grounded multi-agent framework that combines sandboxed Python/PyTorch execution, iterative code repair, citation verification, claim-support auditing, decision control, and structured LaTeX artifact generation. Runtime errors, citation verification failures, and review feedback act as filtering signals to improve execution success, citation validity, local claim support, and workflow completion in controlled evaluations.","AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow  \nAutomation  \nRajesh Kumar, Waqar Ali, Junaid Ahmed, Abdullah Aman Khan, Shaoning Zeng  \narXiv :2607 .02520v1 [ cs .CY] 4 May 2026  \nAbstract—Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify whether generated experiments execute correctly or whether cited sources support generated claims. We present AutoResearch, an execution-grounded multi-agent framework for reliable research workflow automation. AutoResearch couples sandboxed Python/PyTorch execution, iterative code repair, citation verification, claim-support auditing, decision control, and structured LATEX artifact generation. The system treats runtime errors, citation-verification failures, and review-agent feedback as practical filtering signals for generated research artifacts. In controlled evaluations on HumanEval, MBPP, a SciCode subset, citation-validation tasks, claim-support auditing, and small endto-end workflow stress tests, AutoResearch improves execution success, citation validity, local claim support, and workflow completion relative to directly comparable baselines. Code-oriented agents are reported separately as partial comparisons. AutoResearch is intended as a reliability-oriented research assistant, not as a fully autonomous scientist or a standalone manuscript-quality benchmark. (Source Code: AutoResearch)  \nI. INTRODUCTION  \nLLM agents often fail to close the validation loop. We propose AutoResearch (AUTORESEARCH), an execution-grounded multi-agent framework that combines code generation, sandboxed execution, citation validation, and structured researchartifact generation.  \nWe introduce the AutoResearch framework, which treats execution as a hypothesis-validity constraint rather than only a tool for producing outputs. Prior systems such as SWE-agent [24] and AutoGPT [18] use execution within useful agent loops, but they do not explicitly frame runtime evidence as a shared elimination signal across code, citations, and manuscript claims. AUTORESEARCH operationalizes this perspective. Table I positions AUTORESEARCH against prior systems along the axes most relevant to this design choice.  \nRajesh Kumar is with the International Research Center for Complexity Sciences, Hangzhou International Innovation Institute, Beihang University, Hangzhou 311115, China (e-mail: [rajakumarlohano@gmail.com](rajakumarlohano@gmail.com)).  \nWaqar Ali is with the Department of Computer Science, College of Science, Mathematics and Technology, Wenzhou-Kean University, Wenzhou 325060, China ([e-mail: waqar.uestc@yahoo.com](e-mail: waqar.uestc@yahoo.com)).  \nJunaid Ahmed is with Computer Systems Engineering Department, Sukkur IBA University, Sindh, Pakistan (e-mail: [j.bhatti@iba-suk.edu.pk](j.bhatti@iba-suk.edu.pk)).  \nAbdullah Aman Khan is with the International Research Center for Complexity Sciences, Hangzhou International Innovation Institute, Beihang University, Hangzhou 311115, China (e-mail: [rajakumarlohano@gmail.com](rajakumarlohano@gmail.com)).  \nShaoning Zeng is with Yangtze Delta Region Institute (Huzhou), University of Electronic Science and Technology of China, Huzhou 313001, China (e-mail: [zeng@csj.uestc.edu.cn](zeng@csj.uestc.edu.cn)).  \nAUTORESEARCH uses a config-driven, 23-stage pipeline with explicit state transitions (PENDING → RUNNING → FAILED → DONE) . Specialized agents—CodeAgent, BenchmarkAgent, FigureAgent, ReviewAgents—execute code, install dependencies, generate charts, and verify citations. A selfhealing experiment loop detects errors (NaN, crashes) and repairs code. A four-layer citation verification module enforces source validity. MetaClaw stores failure-derived skills and injects them into subsequent runs.  \nOur contributions are threefold:  \n• Systems contribution. We present AUTORESEARCH, a multi-agent research-workflow framework that couples sandboxed experiment exe","cbCaicTsJtBucLeY","https://ap.wps.com/l/cbCaicTsJtBucLeY","pdf",4437983,5,1,15,"English","en",105,"# Introduction\n# Related Work\n## Software engineering agents","[{\"question\":\"AutoResearch如何解决自动化科研代理缺少“验证闭环”的问题？\",\"answer\":\"它把执行结果当作假设有效性的约束条件，将代码生成、沙箱执行、引用校验与结构化文稿生成纳入同一流程，并把运行证据用于排除不可靠产出。\"},{\"question\":\"AutoResearch在可靠性上具体使用了哪些关键模块或机制？\",\"answer\":\"系统耦合了沙箱内Python/PyTorch执行、迭代式代码修复、四层引用验证、声明支持审计，以及基于状态转移（PENDING→RUNNING→FAILED→DONE）的决策控制。\"},{\"question\":\"实验与评估中，AutoResearch主要报告哪些指标，如何理解其局限？\",\"answer\":\"评估关注执行成功、引用有效性、修复成功与端到端流程完成，并强调它不保证科学新颖性或所有文稿声明的正确性；引用校验也只能降低书目幻觉，完整的声明级源支持仍是独立的审计问题。\"}]",1784189615,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"autoresearch-an-execution-grounded-multi-agent-framework-for-reliable-research-workflow-automation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/autoresearch-an-execution-grounded-multi-agent-framework-for-reliable-research-workflow-automation/83664/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"AutoResearch如何解决自动化科研代理缺少“验证闭环”的问题？","Question",{"text":76,"@type":77},"它把执行结果当作假设有效性的约束条件，将代码生成、沙箱执行、引用校验与结构化文稿生成纳入同一流程，并把运行证据用于排除不可靠产出。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"AutoResearch在可靠性上具体使用了哪些关键模块或机制？",{"text":81,"@type":77},"系统耦合了沙箱内Python/PyTorch执行、迭代式代码修复、四层引用验证、声明支持审计，以及基于状态转移（PENDING→RUNNING→FAILED→DONE）的决策控制。",{"name":83,"@type":74,"acceptedAnswer":84},"实验与评估中，AutoResearch主要报告哪些指标，如何理解其局限？",{"text":85,"@type":77},"评估关注执行成功、引用有效性、修复成功与端到端流程完成，并强调它不保证科学新颖性或所有文稿声明的正确性；引用校验也只能降低书目幻觉，完整的声明级源支持仍是独立的审计问题。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]