[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84467-en":3,"doc-seo-84467-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84467,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","A Model-Free Universal AI","General reinforcement learning studies agents in highly general environments with minimal assumptions. This work introduces Universal AI with QInduction (AIQI), a model-free agent that is proven asymptotically ε-optimal and asymptotically ε-Bayes-optimal. Instead of learning policies or explicit environment models, AIQI performs universal induction over distributional action-value return predictors. The paper further extends proof techniques to establish asymptotic ε-optimality of SelfAIXI without ad-hoc assumptions.","A Model-Free Universal AI  \nYegon Kim 1 Juho Lee 1  \n1 Graduate School of AI, KAIST, Seoul, South Korea  \narXiv :2602 .23242v4 [ cs .AI] 13 Jul 2026  \nAbstract  \nIn general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using environment models. This paper introduces Universal AI with QInduction (AIQI), the first model-free agent proven to be asymptotically ε-optimal in general RL. AIQI performs universal induction over distributional action-value functions, instead of policies or environments like previous works. Under a grain of truth condition, we prove that AIQI is strong asymptotically ε-optimal and asymptotically ε -Bayes-optimal. We also apply our novel proof techniques to show asymptotic ε-optimality of SelfAIXI without any ad-hoc assumptions. Our results significantly expand the diversity of known universal agents.  \n1 INTRODUCTION  \nThe theory of general reinforcement learning [GRL; Lattimore, 2014] provides a framework for studying agents in the most general class of environments, those that satisfy only minimal structural assumptions. In this setting, AIXI [Hutter, 2005] is a foundational theoretical model: it combines universal induction [Solomonoff, 1964a,b] with sequential decision theory [Von Neumann and Morgenstern, 1944, Bellman, 1957] to define a Bayes-optimal agent. Although AIXIis uncomputable, it serves as a useful theoretical model of powerful goal-driven AIs [Orseau and Ring, 2012a, Everittet al., 2016], and has also been used to define a universal measure of intelligence, called the Legg-Hutter intelligence [Legg and Hutter, 2007] .  \nInterestingly, AIXI and all other established optimal agents in GRL [Lattimore and Hutter, 2014, Leike et al., 2016a, Cohen et al., 2019, Catt et al., 2023], otherwise known as universal agents or universal AI, are model-based: they ex-  \nplicitly infer and make use of a model of the environment (a world model) . In contrast, the dominant paradigm in practical RL is model-free, learning value functions or policies directly from experience [Watkins, 1989, Williams, 1992, Rummery and Niranjan, 1994, Konda and Tsitsiklis, 1999] . Showing that a model-free algorithm is optimal in GRL has therefore remained elusive and has been repeatedly highlighted as an open problem [Everitt and Hutter, 2018b, Catt, 2022, Hutter et al., 2024] .  \nIn this paper, we present Universal AI with Q-Induction (AIQI), a model-free agent that performs universal induction over return-predictors, objects similar to distributional Q-values [Bellemare et al., 2017] . AIQI is essentially an ε-greedy, on-policy distributional Monte Carlo control algorithm. Under a grain of truth condition [Kalai and Lehrer, 1993], we prove that AIQI achieves strong asymptotic ε -optimality and asymptotic ε-Bayes-optimality in GRL. Our proof techniques can also be used to show asymptotic ε -optimality of Self-AIXI without incurring ad-hoc assumptions such as in the proof by Catt et al. [2023] . Our results significantly expand the class of known universal agents, and provide a blueprint for the analysis of other policy iteration algorithms in general environments.  \n2 BACKGROUND  \n2.1 GENERAL REINFORCEMENT LEARNING  \nWe establish the basic notation for general reinforcement learning. We also provide a summary of all the important symbols in § A. Let A, O, and R ⊆ [0 , 1] denote the finite sets of actions, observations, and rewards, respectively. The percept space is E := O × R, and a history h 1:t ∈ H :=(A × E)∗ is a sequence of actions and percepts. We write h 1:t = a 1 e 1 . . . at et for the history up to time t. Ht :=(A × E)t denotes the set of all histories up to time t. A policy π : H → ∆A maps histories to action distributions, while an environment ν : H×A → ∆E maps history-action pairs to percept distributions. When policy π interacts with  \nAccepted for the 42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026) .  \nenvironment ν, the","cbCaidfs3CqMHtjg","https://ap.wps.com/l/cbCaidfs3CqMHtjg","pdf",6498315,1,24,"English","en",105,"# Abstract\n# Introduction\n# Background\n## General Reinforcement Learning","[{\"question\":\"What problem does the paper address in general reinforcement learning?\",\"answer\":\"It targets the long-standing gap: model-free algorithms have not been shown to be optimal in general reinforcement learning, despite the fact that established universal optimal agents are model-based.\"},{\"question\":\"How does AIQI differ from prior universal agent approaches?\",\"answer\":\"AIQI is model-free and performs universal induction over distributional action-value return predictors, rather than inducing policies or explicit environment/world models.\"},{\"question\":\"What theoretical guarantees does the paper prove for AIQI and SelfAIXI?\",\"answer\":\"Under a grain of truth condition, AIQI is proven to be strong asymptotically ε-optimal and asymptotically ε-Bayes-optimal. The same techniques also show asymptotic ε-optimality of SelfAIXI without ad-hoc assumptions.\"}]",1784195826,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"a-model-free-universal-ai","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/a-model-free-universal-ai/84467/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in general reinforcement learning?","Question",{"text":75,"@type":76},"It targets the long-standing gap: model-free algorithms have not been shown to be optimal in general reinforcement learning, despite the fact that established universal optimal agents are model-based.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does AIQI differ from prior universal agent approaches?",{"text":80,"@type":76},"AIQI is model-free and performs universal induction over distributional action-value return predictors, rather than inducing policies or explicit environment/world models.",{"name":82,"@type":73,"acceptedAnswer":83},"What theoretical guarantees does the paper prove for AIQI and SelfAIXI?",{"text":84,"@type":76},"Under a grain of truth condition, AIQI is proven to be strong asymptotically ε-optimal and asymptotically ε-Bayes-optimal. The same techniques also show asymptotic ε-optimality of SelfAIXI without ad-hoc assumptions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":28,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]