[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86216-en":3,"doc-seo-86216-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86216,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points","Text-based evaluations of Theory of Mind (ToM) in Large Language Models often mirror cognitive tests such as the Sally-Anne task, which can be gamed through pretraining exposure and may not capture functional ToM in naturalistic interaction. The study introduces the Epistemic Asymmetry Schelling Task (EAST), a two-player dialogue benchmark that assesses robust, generalizable social reasoning. Results show a capability gap: only frontier models meet varying epistemic demands, with failures driven by epistemic tracking errors like confusing private and mutual knowledge.","arXiv :2607 . 1 1363v 1 [ cs .CL] 13 Jul 2026  \nBeyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points  \nRoberta Rocca1 , Sami Boukortt1 , Geoff Keeling1,* and Winnie Street1,*  \n1 Google, Paradigms of Intelligence Team, *Joint Last Authors  \nText-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task that can be gamed due to exposure to relevantly similar tasks in pre-training and do not obviously test models’ functional ToM abilities in ways that generalize to naturalistic settings. To address these issues, we introduce the Epistemic Asymmetry Schelling Task (EAST), a two-player dialogue game designed to benchmark robust and generalizable ToM abilities. By requiring LLM-LLM dyads to independently converge on semantic Schelling points under varying states of epistemic transparency, we evaluate whether models can robustly apply ToM to achieve coordination. Our results reveal a significant capability gap in functional social reasoning, with only frontier models successfully navigating the varying epistemic demands of the tasks. Analysis of reasoning traces shows that coordination failures are primarily driven by epistemic tracking errors, such as conflating private knowledge with mutual knowledge. Despite high performance on traditional static benchmarks, our study shows that robust social reasoning and epistemic tracking remain a critical bottleneck, providing concrete targets for future LLM evaluation and development.  \nKeywords: Theory of Mind; LLM Evaluation; Schelling Points; Coordination Game; Epistemic Reasoning  \n1. Introduction  \nTheory of Mind (ToM)– the ability to predict and explain behaviour via the attribution of mental states – is a central determinant of human sociality and cooperation (Premack and Woodruff, 1978; Wimmer and Perner, 1983) . As Large Language Models (LLMs) are increasingly deployed in social settings, acting as tutors, coaches and personal assistants, evaluating their ToM capabilities has become a key research priority (Kosinski, 2024; Sap et al., 2022; Strachan et al., 2024) .  \nWhile LLMs have exhibited strong performance on these tasks (Kosinski, 2024; Street et al., 2025), many evaluation paradigms principally rely on static question-answering tasks, such asthe Sally-Anne false-belief task, adapted from human cognitive psychology (Baron-Cohen et al., 1985; Wimmer and Perner, 1983), as well as narrative-based benchmark suites (Chen et al., 2024; Gandhi et al., 2023; Gu et al., 2024; He et al., 2023; Liu et al., 2025; Sclar et al., 2025; Xu et al., 2024) . Despite their increased scale and complexity, these tasks are subject to non-trivial  \nmethodological concerns (Le et al., 2019; Riemer et al., 2024) . First, LLMs may be able to “game”these tests due to exposure to relevantly similar tasks in pretraining (Shapira et al., 2024; Ullman, 2023) . In particular, it may be that LLMs employ information such as task structure and other linguistic cues as opposed to the relevant mental state information to solve the task. Hence successful performance of the task may not be indicative of ToM having been employed. Second, where ToM inferences have been made, the robustness of the inferences is often left unaddressed. For example, it is unclear whether the ToM inferences employed accurately track the mental states of the target, and whether models are able to coherently use such inferences to make contextually appropriate decisions. For these reasons, it remains unclear whether performance on traditional ToM benchmarks comprehensively assesses the capacity for robust, generalised, interactive social reasoning (Ma et al., 2023), in text-based or embodied settings (Fan et al., 2025; Juneja et al., 2026) .  \nTo address these issues, we introduce the Epis-  \nFigure 1 | Example of the game dynamics for an EAST game in the Asymmetric condition. Prompts are simplified for illustration purposes. See Appen","cbCaisJnV5LFHXFR","https://ap.wps.com/l/cbCaisJnV5LFHXFR","pdf",1759153,2,1,23,"English","en",105,"# Introduction\n## Theory of Mind in LLM evaluation\n## Limitations of static ToM benchmarks\n## Epistemic Asymmetry Schelling Task (EAST)\n## Game setup and epistemic transparency conditions","[{\"question\":\"Why are traditional Sally-Anne-style ToM benchmarks insufficient for evaluating LLMs?\",\"answer\":\"They can be gamed because models may rely on task structure or similar training exposure rather than functional ToM. Robustness of the inferred mental states is also often unclear.\"},{\"question\":\"What is EAST and how does it benchmark Theory of Mind in LLMs?\",\"answer\":\"EAST is a two-player, one-shot text coordination game where LLMs select a shared word based on partner persona and epistemic state without communication. Success requires semantic convergence at appropriate Schelling points.\"},{\"question\":\"What causes coordination failures in EAST according to the analysis?\",\"answer\":\"Coordination failures primarily stem from epistemic tracking errors, such as conflating private knowledge with mutual knowledge.\"}]",1784209529,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-sally-anne-evaluating-theory-of-mind-in-llms-using-epistemic-schelling-points","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-sally-anne-evaluating-theory-of-mind-in-llms-using-epistemic-schelling-points/86216/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are traditional Sally-Anne-style ToM benchmarks insufficient for evaluating LLMs?","Question",{"text":75,"@type":76},"They can be gamed because models may rely on task structure or similar training exposure rather than functional ToM. Robustness of the inferred mental states is also often unclear.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is EAST and how does it benchmark Theory of Mind in LLMs?",{"text":80,"@type":76},"EAST is a two-player, one-shot text coordination game where LLMs select a shared word based on partner persona and epistemic state without communication. Success requires semantic convergence at appropriate Schelling points.",{"name":82,"@type":73,"acceptedAnswer":83},"What causes coordination failures in EAST according to the analysis?",{"text":84,"@type":76},"Coordination failures primarily stem from epistemic tracking errors, such as conflating private knowledge with mutual knowledge.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]