[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81824-en":3,"doc-seo-81824-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81824,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","OPINE-World Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration","Learning environment dynamics from interaction is fundamental for agents that adapt to unfamiliar tasks. Deep-network world models are flexible yet data-hungry and transfer weakly beyond the training distribution. Program-synthesized world models, produced as code by LLMs and refined via counterexample-guided inductive synthesis, improve efficiency and reuse, but prior work rarely scales to pixel-rendered domains with unknown object structure. OPINE-World learns an online object-centric programmatic world model with replay verification and ontology-error-driven exploration, achieving strong results on ARC-AGI-3.","arXiv :2607 .0 153 1v 1 [ cs .AI] 1 Jul 2026  \nOPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration  \nDavid Courtis Wenhao Li Scott Sanner  \nDepartment of Computer Science, University of Toronto [david. courtis@mail. utoronto. ca](david. courtis@mail. utoronto. ca)  \n[chriswenhao. li@mail. utoronto. ca](chriswenhao. li@mail. utoronto. ca)  \n[ssanner@mie. utoronto. ca](ssanner@mie. utoronto. ca)  \nAbstract  \nLearning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned with deep networks are flexible but data-hungry and transfer poorly beyond their training distribution. Program-synthesized world models, written as source code by LLMs and refined through counterexample-guided inductive synthesis (CEGIS), are instead data-efficient and reusable, yet they have been demonstrated mainly on structured-state worlds with a given object vocabulary, and a single program search does not scale to pixel-rendered environments whose object structure must be hypothesized flexibly. We introduce OPINE-World, an LLM agent that learns an object-centric programmatic world model online from interaction. OPINE-World couples two cooperating agents in a loop of hypothesis and test, one acting in the environment and one synthesizing the model in code with replay verification and model-based planning, and it steers exploration with a Bayesian measure of object-type adequacy we call ontology error. We evaluate OPINE-World on ARC-AGI-3, a benchmark for skill-acquisition efficiency in which the object vocabulary, the goal, and the action semantics are withheld. OPINE-World solves 20 of 25 games without per-game training and reaches an action-efficiency score of 78.4 against the human baseline.  \n1 Introduction  \nA world model predicts how an environment changes under each action, and an agent that holds one can plan and transfer to tasks that share a mechanism [Sutton, 1991 , Ha and Schmidhuber, 2018] . In model-based reinforcement learning, world models learned with deep networks are general but data-hungry, and they transfer poorly beyond their training distribution [Hafner et al. , 2020 , Schrittwieser et al. , 2020 , Hafner et al. , 2025 , Kaiser et al. , 2020] . A world model written as a program is the data-efficient alternative, synthesized from logged transitions and refined by counterexample-guided inductive synthesis (CEGIS), and it is inspectable and reusable [Tang et al. , 2024 , Ellis et al. , 2025 , Solar-Lezama et al. , 2006] . Such a program is most economical when it is factored by object type, with one transition rule shared across all objects of a type, so that few parameters support many predictions and a rule learned from one object transfers to the rest [Diuket al. , 2008 , Guestrin et al. , 2003 , Dˇzeroski et al. , 2001] .  \nFactoring helps only when the object partition is approximately correct, and in an open environment that partition is not given and must be inferred from the same interaction the rules are learned from [Kemp et al. , 2006 , Teh et al. , 2006] . Two lines of work mark the extremes of when to commit to it. Program-synthesis-and-plan systems fix the partition at once and plan through the synthesized model, repairing it when no plan reaches reward [Tang et al. , 2024 , Ellis et al. , 2025 ,  \nAhmed et al. , 2025], so a partition fixed before the data identify it carries its error into every rule built on it. Model-free agentic systems fix no structure and act from a natural-language interaction log with no synthesized model [Fox et al. , 2026 , Yao et al. , 2023 , Wang et al. , 2023], keeping no reusable model and no record of what has been identified. The program-synthesis line has been shown on structured-state inputs with a given object vocabulary, and a single program search does not scale to pixel-rendered environments whose object structure must be recovered from observation [Tang et al. , 2024","cbCaijVbZMGySXeE","https://ap.wps.com/l/cbCaijVbZMGySXeE","pdf",1435640,7,1,38,"English","en",105,"# Introduction\n# OPINE-World Method\n## Hypothesis-and-test loop with replay verification\n## Model-based planning and ontology error\n# Evaluation on ARC-AGI-3\n# Contributions","[{\"question\":\"What problem does OPINE-World address in world modeling for agents?\",\"answer\":\"OPINE-World targets data hunger and poor transfer of neural world models, and the scalability limits of program-synthesized models when moving from structured worlds with known object vocabularies to pixel-rendered environments with unknown object structure.\"},{\"question\":\"How does OPINE-World learn its programmatic world model during interaction?\",\"answer\":\"OPINE-World runs a hypothesis-and-test loop with cooperating LLM agents: one proposes dynamics and goals and probes them in the environment, while the other synthesizes code and a critic checks weak generalization using replay verification. A candidate program is accepted only if it reproduces every recorded transition exactly.\"},{\"question\":\"How does OPINE-World guide exploration in unfamiliar object-centered settings?\",\"answer\":\"It uses a Bayesian measure called ontology error to prioritize exploration toward objects whose behaviors the current object types do not yet explain, helping the agent discover an object ontology from interaction.\"}]","OPINE-World Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration | PDF",1784176394,96,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"opine-world-programmatic-world-modeling-with-ontology-error-prioritized-interactive-exploration","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/opine-world-programmatic-world-modeling-with-ontology-error-prioritized-interactive-exploration/81824/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does OPINE-World address in world modeling for agents?","Question",{"text":77,"@type":78},"OPINE-World targets data hunger and poor transfer of neural world models, and the scalability limits of program-synthesized models when moving from structured worlds with known object vocabularies to pixel-rendered environments with unknown object structure.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does OPINE-World learn its programmatic world model during interaction?",{"text":82,"@type":78},"OPINE-World runs a hypothesis-and-test loop with cooperating LLM agents: one proposes dynamics and goals and probes them in the environment, while the other synthesizes code and a critic checks weak generalization using replay verification. A candidate program is accepted only if it reproduces every recorded transition exactly.",{"name":84,"@type":75,"acceptedAnswer":85},"How does OPINE-World guide exploration in unfamiliar object-centered settings?",{"text":86,"@type":78},"It uses a Bayesian measure called ontology error to prioritize exploration toward objects whose behaviors the current object types do not yet explain, helping the agent discover an object ontology from interaction.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,117,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":113,"doc_module":4,"doc_module_name":47,"category_name":114,"show_sort_weight":115,"slug":116},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]