[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83596-en":3,"doc-seo-83596-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83596,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games","We investigate learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games, where agents benefit from reasoning about opponents’ stochastic policies and their own. The work delivers three contributions: dataset construction for policies in a given game, representation-learning methods for policy embeddings, and downstream evaluations to measure effectiveness. Experiments evaluate dataset methods, embedding methods, and tasks on Kuhn and Leduc Poker, showing that the learned embeddings contain meaningful behavioral structure.","arXiv :2607 .0 1498v 1 [ cs .LG] 1 Jul 2026  \nTowards Learning Representations of Policies in Two-Player Zero-Sum  \nImperfect-Information Games  \nKevin Wang∗ KEVIN A WANG @BROWN . EDU  \nKevin Yang∗ KEVIN   C YANG @BROWN . EDU  \nArjun Prakash ARJUN PRAKASH @BROWN . EDU  \nAmy Greenwald AMY   GREENWALD @BROWN . EDU  \nBrown University  \nAbstract  \nWe investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introduce methods of creating datasets of policies for a given game. Second, we propose methods to learn policy representations. Third, we introduce downstream tasks to evaluate the effectiveness of such representations.  \nWe evaluate each dataset method, embedding method, and downstream task on Kuhn and Leduc Poker. Although our methods are very basic, we demonstrate that useful behavioral representations are present in the learned embeddings. To our knowledge, this work is among the first to systematically compare self-supervised learning techniques for learning policy representations in games. Our code is available at [https://github.com/VitamintK/ssl-project](https://github.com/VitamintK/ssl-project) for others to extend.  \n1. Introduction  \nIn two-player zero-sum imperfect-information games, agents can benefit from reasoning about their opponent’s policies, and their own. In general, policies in such games are stochastic.1  \nFor example, an agent seeking to play the Nash equilibrium at some decision point in a hand of poker may perform lookahead search by imagining that each player plays some policy for a few steps, and then evaluating the resulting public belief state. This depth-limited search is typically done tabularly2 [2, 21], which is feasible since such a depth-limited policy has a tractable size. However, in games with larger public belief states, such a policy becomes intractable to enumerate. Thus, to perform a version of such a search, an agent will need to reason with compact representations of policies.  \nSo we want to learn good, compact representations of policies. However, there is little existing research towards this end, especially in the context of two-player zero-sum imperfect-information games. There are few existing methods to learn policy embeddings, or comparisons between methods. In order to compare methods, we also need an extensive set of evaluations, which does not exist.  \nIn this work, we begin an extensive, systematic approach to the problem of learning compact, useful representations of policies, particularly in two-player zero-sum imperfect-information games.  \n*  \nEqual contribution.  \n1. In contrast, in single-agent or perfect-information settings, reasoning about individual actions typically suffices – e.g. lookahead search in chess and go.  \n2. e.g. using counterfactual regret minimization or linear programming  \n© K. Wang∗ , K. Yang∗ , A. Prakash & A. Greenwald.  \nTOWARDS LEARNING REPRESENTATIONS OF POLICIES IN GAMES  \nWe are particularly interested in methods that allow decoding embeddings into policies and in representations that can predict future payoffs.  \nWe do this in three parts:  \n1. We introduce three methods of creating datasets of policies for a given game. (Section 3.1)  \n2. We propose several methods of varying complexity for learning representations of policies. We also reimplement an existing method. (Section 3.2)  \n3. We introduce several downstream tasks to evaluate the usefulness of the policy representations, and we use them to evaluate the methods. (Section 3.3)  \n2. Preliminaries  \nImperfect Information Games An imperfect-information game (IIG) is one in which each player may not know the true state of the world. A two-player zero-sum game is one in which there are two players, and the players’ payoffs sum to 0 . Formally, an IIG is given by the tuple [19]:⟨S, A, O, I, R, T, O , C, tmax ⟩ , where S is the space of game states, A is the action space, O ","cbCaicYQgPh4xakU","https://ap.wps.com/l/cbCaicYQgPh4xakU","pdf",3157789,2,1,22,"English","en",105,"# Introduction\n# Preliminaries\n## Imperfect Information Games\n## Related Works\n# Methods\n## Methods for Making Datasets of Policies\n## Methods for Learning Embeddings\n# Downstream Tasks","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It addresses how to learn compact, useful policy representations (embeddings) for two-player zero-sum imperfect-information games.\"},{\"question\":\"What are the three main contributions?\",\"answer\":\"The paper (1) introduces dataset-creation methods for policies in a given game, (2) proposes methods to learn policy embeddings, and (3) defines downstream tasks to evaluate embedding effectiveness.\"},{\"question\":\"How are the methods evaluated and on which games?\",\"answer\":\"Datasets, embedding methods, and downstream tasks are evaluated on Kuhn and Leduc Poker to test whether the embeddings capture useful behavioral representations.\"}]",1784189087,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-learning-representations-of-policies-in-two-player-zero-sum-imperfect-information-games","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/towards-learning-representations-of-policies-in-two-player-zero-sum-imperfect-information-games/83596/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"It addresses how to learn compact, useful policy representations (embeddings) for two-player zero-sum imperfect-information games.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the three main contributions?",{"text":80,"@type":76},"The paper (1) introduces dataset-creation methods for policies in a given game, (2) proposes methods to learn policy embeddings, and (3) defines downstream tasks to evaluate embedding effectiveness.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the methods evaluated and on which games?",{"text":84,"@type":76},"Datasets, embedding methods, and downstream tasks are evaluated on Kuhn and Leduc Poker to test whether the embeddings capture useful behavioral representations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]