[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86177-en":3,"doc-seo-86177-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86177,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory","LLM agents face a dilemma: static safety prompts limit exploration, while free tool use and multi-agent debate often cause sandbox violations. This work splits responsibilities across specialized roles: a Disrupter proposes unconventional actions, a Validator enforces hard runtime checks at the tool gateway, and a Broker imports distant yet relevant analogies. Execution failures are compiled via MCTS into signed constraint patches (“Scars”) cached for cohort inheritance. In a spatial-semantic sandbox, Scars prevent all breaches and reduce token use by 15.1%, while CAS bandwidth control cuts overall token costs by 55.9% under constraints.","Heterogeneous Agent Cohorts for Safe Open-Ended  \nExploration  \nwith Runtime Constraint Memory  \nLIU TENGJIAO  \nFounder & Researcher, [psi.run](psi.run), [psi@psi.run](psi@psi.run)  \narXiv :2607 . 1 1226v 1 [ cs .AI] 13 Jul 2026  \nAbstract—LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free rein with tools and multi-agent debate, and safety violations quickly follow. Rather than forcing a single model to juggle both creativity and caution, we separate the concerns across specialized roles. A Disrupter generates unconventional proposals, a Validator enforces hard runtime checks at the tool gateway, and a Broker pulls in distant but relevant analogies. Failures are not discarded—they are compiled, via MCTS, into compact, signed constraint patches we call Scars. These patches are cached locally and inherited by future cohorts, turning repeated failures into reusable, low-cost runtime constraints. In a spatial-semantic sandbox (N=20 runs, p¡0.01), our cohort reaches remote targets where debate fails, the Validator prevents all executed breaches, and Scars reduce token consumption by 15.1% by avoiding redundant validator checks. Furthermore, credit-based Communication Allocation Scores (CAS) restrict outbound bandwidth, reducing overall token costs by 55.9% under resource constraints.  \nI. INTRODUCTION  \nExploring open-ended environments and generating nontrivial discoveries (such as formal mathematical proofs, algorithmic synthesis, or novel molecular designs) remains a defining goal in artificial intelligence. Recently, with the progress of Large Language Models (LLMs) in reasoning and tool-use, language agents have transitioned into automating research workflows. For instance, Google DeepMind’s Co-Scientist [12] coordinates multi-agent networks for literature review, while Sakana AI’s The AI Scientist [10], [11] automates scientific writing and code evolution. However, in open-ended exploration, agent systems run into a fundamental conflict between creative novelty and runtime safety. Discovering non-trivial serendipitous solutions usually requires executing high-entropy proposals, which in turn risks triggering severe sandbox failures. Consider an LLM agent deployed to automate scientific discovery, asked to design a new chemical synthesis experiment. An unrestricted agent might attempt to call tools that execute unsafe heating cycles or write to forbidden directories, damaging the host environment. Conversely, an agent locked behind static safety prompts may refuse to suggest any unconventional reaction conditions, failing to discover novel chemical pathways. This trade-off between creative exploration and system safety represents a major bottleneck for autonomous agents.  \nExisting frameworks for agentic scientific exploration present clear limitations. Reflexion [2] forces the model to self-correct, only to get trapped inside its own parametric prior. MemGPT [1] treats memory as a paging system, accumulating context clutter and driving tool parameter drift [7] . AutoGen [19] enables multi-agent conversations, but lacks safety sandboxes—leading agents to either self-censor into sterility or delete directories in high-entropy states.  \nThis paper extends the single-agent Scar memory mechanism of [5] to a three-role cohort. The key intuition is that role specialization redistributes the exploration-safety trade-off: the Disrupter generates high-entropy proposals that a solitary agent would self-censor, while the Validator and Broker contribute targeted safety and knowledge functions without diluting eachother’s objectives.  \nWe separate exploration and safety control into distinct, specialized components. Under this framework, the human operator defines the initial cohort topology and validation objectives, while the agent cohort runs autonomously in the sandbox. The cohort generates out-of-distribution (OOD) discoveries throug","cbCaifChDw2B3Sqm","https://ap.wps.com/l/cbCaifChDw2B3Sqm","pdf",378593,4,1,12,"English","en",105,"# Introduction\n# Research Questions and Hypotheses\n# Related Work","[{\"question\":\"Why do current LLM agent approaches struggle with open-ended exploration?\",\"answer\":\"Static safety instructions tend to prevent novel, high-entropy proposals, while granting unrestricted tool access or debate can quickly lead to sandbox violations and unsafe actions.\"},{\"question\":\"How does the proposed heterogeneous cohort improve the exploration–safety trade-off?\",\"answer\":\"It assigns distinct roles: the Disrupter generates unconventional proposals, the Validator performs runtime safety gating at the tool gateway, and the Broker supplies relevant cross-domain analogies without diluting other objectives.\"},{\"question\":\"What are “Scars” and how do they affect runtime efficiency?\",\"answer\":\"Scars are cryptographically signed constraint patches generated by compiling execution failures with MCTS. They are cached locally and inherited by future cohorts, reducing redundant validation and lowering token consumption.\"}]",1784209131,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"heterogeneous-agent-cohorts-for-safe-open-ended-exploration-with-runtime-constraint-memory","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/heterogeneous-agent-cohorts-for-safe-open-ended-exploration-with-runtime-constraint-memory/86177/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do current LLM agent approaches struggle with open-ended exploration?","Question",{"text":75,"@type":76},"Static safety instructions tend to prevent novel, high-entropy proposals, while granting unrestricted tool access or debate can quickly lead to sandbox violations and unsafe actions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed heterogeneous cohort improve the exploration–safety trade-off?",{"text":80,"@type":76},"It assigns distinct roles: the Disrupter generates unconventional proposals, the Validator performs runtime safety gating at the tool gateway, and the Broker supplies relevant cross-domain analogies without diluting other objectives.",{"name":82,"@type":73,"acceptedAnswer":83},"What are “Scars” and how do they affect runtime efficiency?",{"text":84,"@type":76},"Scars are cryptographically signed constraint patches generated by compiling execution failures with MCTS. They are cached locally and inherited by future cohorts, reducing redundant validation and lowering token consumption.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]