[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82922-en":3,"doc-seo-82922-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82922,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization","Compositional generalization—the ability to understand and produce novel combinations of known components—remains a core challenge in modern AI. ClassicLogic introduces a benchmark suite built from four classic logic puzzles: Sudoku, KenKen, Kakuro, and Futoshiki. Each game includes a hierarchical, explicit knowledge base where multi-step solving strategies are formally composed from simpler foundational strategies. This enables fine-grained evaluation of reasoning from basic rule learning to mathematically validated, increasing difficulty, supporting neuro-symbolic and advanced AI systems. ","ClassicLogic: A KNOWLEDGE-DRIVEN BENCHMARK OF CLASSIC PUZZLE GAMES FOR EVALUATING COMPOSITIONAL  \nGENERALIZATION  \nA PREPRINT  \narXiv :2607 .05 185v 1 [ cs .AI] 6 Jul 2026  \n Mahnoor Shahid  \nUniversität Duisburg-Essen Germany  \n[mahnoor.shahid@uni-due.de](mahnoor.shahid@uni-due.de)  \n Hannes Rothe  \nUniversität Duisburg-Essen Germany  \n[hannes.rothe@uni-due.de](hannes.rothe@uni-due.de)  \nABSTRACT  \nCompositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge for modern artificial intelligence. While few benchmarks exist, many focus on linguistic tasks and lack complex, explicit compositional structures.  \nWe introduce ClassicLogic, a new benchmark suite designed to evaluate an agent’s ability to learn and compose problem-solving strategies. The benchmark consists of four classic logic puzzles: Sudoku, KenKen, Kakuro, and Futoshiki. Its core innovation is a hierarchical, explicit knowledge base for each game, where complex solving strategies are formally defined as compositions of simpler, foundational strategies. This structure allows for fine-grained evaluation of an agent’s reasoning capabilities, from learning basic rules to applying multi-step compositional strategies to solve puzzles of increasing, mathematically validated difficulty. The open-source benchmark provides a challenging new testbed for advancing neuro-symbolic and  \nother advanced AI reasoning systems. The benchmark is open-source and available at: [https:](https:)//[github.com/Place-Beyond-Bytes/classic_games_benchmark.git](github.com/Place-Beyond-Bytes/classic_games_benchmark.git).  \nKeywords Compositional Generalization, Benchmark, Neuro-Symbolic AI, Logic Puzzles, Multi-Step Reasoning  \n1 Introduction  \nThe ability to flexibly combine existing knowledge to solve new problems is a hallmark of human intelligence (Sternberg, 1984; Cosmides and Tooby, 1997) . This capacity, termed as compositional generalization Keysers et al.(2019); Wiedemer et al. (2023), enables us to generate a potentially infinite range of complex ideas and behaviors from a finite set of known components (Fodor and Pylyshyn, 1988) . Modern artificial intelligence, particularly large-scale models, has demonstrated remarkable capabilities in processing and generating human-like data(Zhao et al., 2023) . Yet, a critical gap persists between their pattern-matching prowess and the robust, systematic reasoning characteristic of human intelligence(Marcus, 2018) . These models often fail at tasks requiring logical deduction and multi-step planning, revealing a fundamental weakness in their ability to generalize systematically (Patil and Jadon, 2025; Cao et al., 2025) .  \nTo address this, the community has developed benchmarks such as SCAN (Lake and Baroni, 2018) and COGS (Kim and Linzen, 2020) to test for compositional generalization, primarily in the domain of natural language processing. However, a gap exists for benchmarks that test compositional reasoning in more structured, symbolic problem-solving domains (Wang et al., 2024) . Such domains require not only recognizing patterns but also building and executing explicit, multi-step strategies.  \nIn this paper, we introduce ClassicLogic, a novel benchmark designed to provide a transparent and challenging testbed for compositional reasoning. Our contribution is a suite of four classic logic puzzles—Sudoku, KenKen, Kakuro, and Futoshiki—that are procedurally generated and comes with a hierarchical knowledge base (KB) of solving strategies, where complex solving strategies are explicitly defined as compositions of simpler, atomic rules. For instance, advanced strategies are explicitly defined as compositions of more fundamental ones (e.g., identifying a ‘naked_pair’ in Sudoku  \nis a composition of the ‘naked_single’ and ‘constraint_propagation’ strategies) . This structure allows us to dissect the reasoning process and evaluate three distinct and crucial forms of comp","cbCaish0WiDYy4yn","https://ap.wps.com/l/cbCaish0WiDYy4yn","pdf",3241992,3,1,13,"English","en",105,"# Introduction\n## Benchmarks for Compositional Generalization\n# Related Work","[{\"question\":\"What problem does ClassicLogic address in AI research?\",\"answer\":\"ClassicLogic targets compositional generalization: enabling AI systems to combine known components into novel, multi-step reasoning for new puzzle instances.\"},{\"question\":\"Which classic logic games are included in the ClassicLogic benchmark?\",\"answer\":\"The benchmark contains four puzzles: Sudoku, KenKen, Kakuro, and Futoshiki.\"},{\"question\":\"How does the benchmark represent solving strategies and difficulty?\",\"answer\":\"It provides a hierarchical, explicit knowledge base where complex strategies are defined as compositions of simpler atomic strategies, supporting fine-grained assessment across increasing, validated difficulty levels.\"}]",1784183968,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"classiclogic-a-knowledge-driven-benchmark-of-classic-puzzle-games-for-evaluating-compositional-generalization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/classiclogic-a-knowledge-driven-benchmark-of-classic-puzzle-games-for-evaluating-compositional-generalization/82922/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ClassicLogic address in AI research?","Question",{"text":75,"@type":76},"ClassicLogic targets compositional generalization: enabling AI systems to combine known components into novel, multi-step reasoning for new puzzle instances.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which classic logic games are included in the ClassicLogic benchmark?",{"text":80,"@type":76},"The benchmark contains four puzzles: Sudoku, KenKen, Kakuro, and Futoshiki.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the benchmark represent solving strategies and difficulty?",{"text":84,"@type":76},"It provides a hierarchical, explicit knowledge base where complex strategies are defined as compositions of simpler atomic strategies, supporting fine-grained assessment across increasing, validated difficulty levels.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]