[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81779-en":3,"doc-seo-81779-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81779,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","AGI Maze as a Benchmark Framework for World Modeling Agents","Large language models function as next-token predictors, yet this default mode does not reliably yield persistent, manipulable representations of an external world. Many “reasoning” tasks become harder with partial observability, statefulness, memory requirements, and hidden-state hypotheses. AGI Maze provides a lightweight, grid-based, stateful environment family with clean APIs and multiple difficulty regimes. It evaluates LLMs and introduces a memory-augmented baseline, showing that even small mazes remain unsolved within practical step budgets.","AGI Maze as a Benchmark Framework for World  \nModeling Agents  \nAlexey Potapov  \nSingularityNET Foundation  \n[alexey@singularitynet.io](alexey@singularitynet.io)  \nAbstract  \nLarge language models (LLMs) are powerful pattern-completion systems, but their default operating mode – predicting the next token from a static context – does not reliably produce persistent, manipulable representations ofan external world. Many tasks that look like \"reasoning\"in text become substantially harder once the environment is partially observable, stateful, and requires memory and structured hypotheses about hidden state.  \nAGI Maze is a lightweight framework for building such environments without requiring highdimensional sensory inputs. It provides a family of grid-based maze tasks with a clean API and multiple difficulty regimes. The goal is to create benchmarks where agents must learn and use world state representations, not just infer a local rule over readily provided observations.  \nWe provide an initial evaluation of several vanilla LLMs on simple mazes showing that they fail to represent mazes internally at LLM inference time. We also introduce a baseline agent, which is allowed to use its message history as a working memory to construct descriptions of observations at agentic runtime. Although this can improve performance, it is still insufficient for an LLM agent to reliably solve even small mazes within a step budget that is more than enough for humans.  \n1. Motivation: from LLM limitations to world models  \n1.1 LLMs are static predictors, not persistent world simulators  \nA core limitation of current LLM-only agents is that the model itself is static: the only \"dynamics\"typically come from:  \n• adding text to the prompt (message history as working memory),  \n• retrieving external text (RAG) into the prompt (long-term memory),  \n• and feeding tool outputs back as text.  \nThis creates two intertwined problems:  \n1. Memory as text is inefficient. Some tasks require explicit state tracking, revisiting earlier observations, and performing multi-step search. Encoding everything as unstructured text leads to brittle behavior and high token cost.  \n2. Representation is not guaranteed. Next-token prediction encourages extracting whatever information is needed to produce the next output, not maintaining a stable, manipulable  \nrepresentation of “what is true in the world.” In language tasks this deficiency is partially hidden because language already is a world description; in interactive environments, it becomes obvious.  \nWe discussed how the architecture and the training objective of LLMs are connected with their practical limitation in the paper [1] in detail.  \nA complementary angle comes from empirical observations that internal activations can encode much of the final output early in the network, which supports the view that many models are best described as efficient predictors rather than explicit world-state modelers. See e.g. [2] (the discussion around intermediate-layer information content including possibility to remove, swap and iterate over blocks of transformers in LLMs, including references [3-6]) .  \n1.2 Why “world models” need more than rule discovery  \nIn the language domain, \"world modeling\" is often reduced to \"discover the rule of a game.\" But for agents, a full notion of world modeling also includes:  \n• representing latent state (what you cannot currently observe),  \n• maintaining beliefs under uncertainty,  \n• updating those beliefs with new evidence,  \n• reasoning over representations (maps, graphs, causal schemas),  \n• and using memory efficiently (working memory, episodic memory, long-term knowledge) .  \nIn other words: world models are about state and representation, not only about rules.  \n1.3 World modeling testbeds and their limitations  \nAGI-ARC-3 [7] is an important benchmark for generalization across tasks and for inferring hidden rules. However, many ARC-like settings:  \n• are effectively fully observable","cbCaivd6uigxTMsV","https://ap.wps.com/l/cbCaivd6uigxTMsV","pdf",354972,4,1,10,"English","en",105,"# Abstract\n# Motivation: from LLM limitations to world models\n## LLMs are static predictors, not persistent world simulators\n## Why “world models” need more than rule discovery\n## World modeling testbeds and their limitations\n# AGI Maze: a lightweight testbed for state, memory, and representation\n## Design goals","[{\"question\":\"What problem does AGI Maze aim to measure in AI agents?\",\"answer\":\"AGI Maze benchmarks whether agents can learn and use world state representations, not merely infer local rules from observations.\"},{\"question\":\"Why are standard LLM-based agents insufficient for world modeling?\",\"answer\":\"LLMs operate as static next-token predictors, and representing state as unstructured text is inefficient while also failing to guarantee stable, manipulable world representations under interaction.\"},{\"question\":\"How does AGI Maze differ from other ARC-like benchmarks?\",\"answer\":\"It emphasizes partial observability, persistent queryable state across time, localization/mapping, and long-horizon memory demands, which ARC-like tasks often do not strongly test.\"}]",1784176092,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"agi-maze-as-a-benchmark-framework-for-world-modeling-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/agi-maze-as-a-benchmark-framework-for-world-modeling-agents/81779/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does AGI Maze aim to measure in AI agents?","Question",{"text":75,"@type":76},"AGI Maze benchmarks whether agents can learn and use world state representations, not merely infer local rules from observations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are standard LLM-based agents insufficient for world modeling?",{"text":80,"@type":76},"LLMs operate as static next-token predictors, and representing state as unstructured text is inefficient while also failing to guarantee stable, manipulable world representations under interaction.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AGI Maze differ from other ARC-like benchmarks?",{"text":84,"@type":76},"It emphasizes partial observability, persistent queryable state across time, localization/mapping, and long-horizon memory demands, which ARC-like tasks often do not strongly test.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]