[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81838-en":3,"doc-seo-81838-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81838,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Coding agents can replicate scientific machine learning papers","Scientific machine learning papers often make precise computational claims, yet prompts that ask an agent to reproduce them do not reliably preserve execution progress or validate whether the generated evidence truly supports the claims. A Paper-replication workflow is introduced to map each selected paper claim to a target with recorded evidence, reconstruct the method, run computational experiments, and link outputs to provenance and claim-level comparisons. Validation checks must pass before completion, and evaluation across multiple runs confirms workspace evidence governs success.","Coding-agents can replicate scientific machine learning papers  \nAtharva Hansa,1,∗ and Ilias Bilionisa  \na School of Mechanical Engineering, Purdue University, West Lafayette, IN, 47907, USA  \nAbstract  \nScientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that the 95% predictive credible interval covers the test data. A coding agent can be prompted to replicate those claims from paper materials alone, but the prompt does not by itself reliably preserve progress or check whether generated evidence supports the paper’s claims. We introduce Paper-replication, a workflow that makes each selected paper claim a target with recorded evidence, and implement it as a coding-agent skill. The workflow makes the agent record those targets, reconstruct the paper’s method, run computational experiments, link generated outputs to provenance and comparisons with the paper’s claims, record where matched evidence appears in the replication report, and pass validation checks before completion. We evaluate Paper-replication on twelve independent runs across four scientific machine learning papers. All twelve workspaces pass the completion gate, and all 158 recorded targets are matched with report coverage. Even in this completed workspace state, repeated runs differ in how papers are divided into targets, in numerical fidelity to the source papers, in elapsed replication time, in the number of intermediate executions replaced before final evidence is accepted, and in the rules used to accept evidence. Paper-replication makes completion depend on workspace evidence and validation checks rather than on the agent’s final message.  \nKeywords: Coding agents, Harness engineering, Paper replication, Scientific reproducibility, Scientific machine learning  \nJul 2026  \nsolvers and Rackauckas  \nmachine learning papers neural differential equaet al., 2020; Hans et al.,  \n∗ Corresponding author.  \nEmail address: [atharva.hans@lilly.com](atharva.hans@lilly.com) (Atharva Hans)  \n1Present address: Eli Lilly and Company, Indianapolis, IN 46285, USA  \nrecording evidence for each claim. Replicating a paper from paper materials alone is a narrower and harder task than rerunning a released package. The available materials may include the LATEX source, figures, tables, appendices, and data references, but they need not include sampled training sets, seeds, optimizer states, sampler states, preprocessing conventions, or plotting rules. In machine learning, reproducibility taxonomies treat a paper-only record as weaker evidence than a runnable package with code, data, dependencies, and environment information (Tatman et al., 2018) . Surveys of artificial-intelligence papers also find that empirical details are often underreported (Gundersen and Kjensmo, 2018) . Stochastic training makes these omissions matter because data sampling, initialization, implementation choices, and hyperparameters can change the result even when the algorithmic description stays fixed (Henderson et al., 2018; Bouthillier et al., 2021; Pineau et al., 2021) .  \nResearch code generation methods address one part of paper replication: they convert paper descriptions into code or implementation tasks. PaperCoder is a multi-agent framework that turns machine learning papers into code repositories through planning, implementation analysis, and modular code generation (Seo et al., 2025) . ResearchCodeBench converts recent machine learning papers into executable implementation challenges (Hua et al., 2026), and ResearchCodeAgent uses a multiagent system with planning, memory, and actions to codify methods from the machine learning literature (Gandhi et al., 2025) . Other benchmarks evaluate algorithmic reproduction from natural-language-processing papers (Xiang et al., 2025), reconstruction of masked language-modeling research code (Yanet al., 2025), and progressive code-masking experiments (Kim  \net al., 2025) . These stud","cbCaijlxFpdILqBk","https://ap.wps.com/l/cbCaijlxFpdILqBk","pdf",317147,7,1,16,"English","en",105,"# Abstract\n## Paper-replication workflow\n## Evidence recording and validation\n## Evaluation results\n## Related work: code generation and reproduction benchmarks","[{\"question\":\"Why is replicating scientific machine learning paper claims from paper materials alone difficult for coding agents?\",\"answer\":\"Paper-only prompts may fail to preserve progress and do not ensure that generated evidence genuinely supports the paper’s claims. Missing implementation and experimental details make validation non-trivial.\"},{\"question\":\"What does the Paper-replication workflow require before an agent can finish?\",\"answer\":\"It records claim targets with evidence, reconstructs the paper’s method, runs computational experiments, links outputs to provenance and claim comparisons, and completes only after validation checks pass.\"},{\"question\":\"How was Paper-replication evaluated, and what was observed across repeated runs?\",\"answer\":\"The workflow was tested in twelve independent runs across four papers. All workspaces completed with all recorded targets matched, while repeated runs differed in target partitioning, numerical fidelity, replication time, intermediate execution handling, and evidence-acceptance rules.\"}]","Coding agents can replicate scientific machine learning papers | PDF",1784176554,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"coding-agents-can-replicate-scientific-machine-learning-papers","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/coding-agents-can-replicate-scientific-machine-learning-papers/81838/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-01","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is replicating scientific machine learning paper claims from paper materials alone difficult for coding agents?","Question",{"text":77,"@type":78},"Paper-only prompts may fail to preserve progress and do not ensure that generated evidence genuinely supports the paper’s claims. Missing implementation and experimental details make validation non-trivial.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What does the Paper-replication workflow require before an agent can finish?",{"text":82,"@type":78},"It records claim targets with evidence, reconstructs the paper’s method, runs computational experiments, links outputs to provenance and claim comparisons, and completes only after validation checks pass.",{"name":84,"@type":75,"acceptedAnswer":85},"How was Paper-replication evaluated, and what was observed across repeated runs?",{"text":86,"@type":78},"The workflow was tested in twelve independent runs across four papers. All workspaces completed with all recorded targets matched, while repeated runs differed in target partitioning, numerical fidelity, replication time, intermediate execution handling, and evidence-acceptance rules.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":108,"slug":138},19,"General","general"]