[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82248-en":3,"doc-seo-82248-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82248,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents","Large language model (LLM) agents are increasingly used for AI-driven scientific discovery, where they repeatedly propose hypotheses, test them, and update beliefs using evidence. Current systems hide these hypothesis–test–evidence–belief steps in unstructured logs, preventing both agent- and human-level auditing. The Hypothesis Evolution Protocol (HEP) makes the cycle explicit and traceable: hypotheses are stored as persistent auditable objects, belief probabilities change only via attached evidence, and hypotheses follow a lifecycle from proposal to resolution.","arXiv :2607 .09 195v 1 [ cs .AI] 10 Jul 2026  \nToward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents  \nIzumi Takahara1, * and Teruyasu Mizoguchi1,†  \n1 Institute of Industrial Science, The University of Tokyo, Tokyo, Japan 153-8505  \n*[kougen@iis.u-tokyo.ac.jp](kougen@iis.u-tokyo.ac.jp), †[teru@iis.u-tokyo.ac.jp](teru@iis.u-tokyo.ac.jp)  \nLarge language model (LLM) agents are increasingly expected to play a central role in AIdriven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis–test–evidence–belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.  \n1. Introduction  \nThe rapid progress of large language models (LLMs), and in particular of LLM-based agents, has attracted intense interest as a potential driver of scientific discovery across a wide range of domains [1–4] . Modern LLMs combine broad scientific knowledge with flexible reasoning capabilities, and when equipped with tools [5–7], skills [8, 9], memory [10, 11], and reasoning strategies such as planning and reflection [5, 12, 13], the resulting agents have the potential to drive scientific research flexibly and autonomously.  \nTo date, LLM agents have been developed toward the autonomous execution of research in several fields. Prominent examples include the end-to-end automation of artificial intelligence research and engineering [1, 14] and the automation of biomedical discovery [2, 3, 15, 16], as well as chemistry agents that couple LLMs with domain tools and robotic laboratories [17–19] . In materials science, LLM agents have been actively applied both to computational materials design [20–26] and to the automation of materials synthesis and characterization [27–31], and they are emerging as a practical route to accelerating the progress of materials research [4] .  \nThe term “autonomous research” by LLM agents, however, encompasses tasks with substantially different levels of agency [32] . Much of the work above automates predefined workflows, in which the agent reliably executes a prescribed sequence of computations, syntheses, or measurements [17–19, 26], or performs optimization toward an externally specified goal [27, 28] . Atask demanding substantially higher agency is to answer the question of why: to arrive at explanatory understanding that is earned through repeated cycles of generating mechanistic hypotheses, devising the empirical tests that can discriminate among them, and revising beliefs against the resulting evidence. Steps in this direction include agents that generate and rank scientific hypotheses [2, 16, 33–36], frameworks that place falsification at the center of the research process [37, 38], belief-guided open-ended exploration [39], and closed-loop systems that iterativelyrefine hypotheses against data, experiments, or simulations [3, 14, 15, 40, 41] .  \nNevertheless, in existing agents, neither the agent itself nor the human researcher can readily trace which hypotheses the agent entertained, which tests it performed, and how the resulting evidence moved its beliefs. Hypotheses and beliefs typically remain implicit in free-text reasoning or i","cbCaigcPmjuZpVvQ","https://ap.wps.com/l/cbCaigcPmjuZpVvQ","pdf",1345177,1,17,"English","en",105,"# Introduction\n## Motivation for auditable scientific reasoning\n## Limits of current LLM agent traces\n## Hypothesis Evolution Protocol (HEP)\n## Evaluation on materials-science tasks","[{\"question\":\"Why do existing LLM agents fail to support auditing of scientific reasoning?\",\"answer\":\"Because hypotheses, tests, and belief updates are typically buried in unstructured logs or implicit internal states, leaving researchers unable to trace how evidence changed beliefs.\"},{\"question\":\"What is the Hypothesis Evolution Protocol (HEP)?\",\"answer\":\"HEP turns the proposing–testing–revising process into an explicit, traceable protocol where each hypothesis is an auditable persistent object with belief probability governed by attached evidence.\"},{\"question\":\"How does HEP differ from planning-style agent behavior?\",\"answer\":\"HEP provides a disciplined hypothesis–test–evidence–belief cycle, and experiments on materials-science tasks show that an HEP-equipped agent operates this cycle in ways planning-style agents lack, with deeper protocol usage scaling with base-LLM capability.\"}]",1784179149,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"toward-auditable-ai-scientists-a-hypothesis-evolution-protocol-for-llm-agents","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/toward-auditable-ai-scientists-a-hypothesis-evolution-protocol-for-llm-agents/82248/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing LLM agents fail to support auditing of scientific reasoning?","Question",{"text":75,"@type":76},"Because hypotheses, tests, and belief updates are typically buried in unstructured logs or implicit internal states, leaving researchers unable to trace how evidence changed beliefs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Hypothesis Evolution Protocol (HEP)?",{"text":80,"@type":76},"HEP turns the proposing–testing–revising process into an explicit, traceable protocol where each hypothesis is an auditable persistent object with belief probability governed by attached evidence.",{"name":82,"@type":73,"acceptedAnswer":83},"How does HEP differ from planning-style agent behavior?",{"text":84,"@type":76},"HEP provides a disciplined hypothesis–test–evidence–belief cycle, and experiments on materials-science tasks show that an HEP-equipped agent operates this cycle in ways planning-style agents lack, with deeper protocol usage scaling with base-LLM capability.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]