[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85992-en":3,"doc-seo-85992-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85992,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","MafiaScope: 非侵入式、时间分辨的信念探测，用于社交推理游戏中的LLM智能体","LLM agents’ public actions provide limited insight into their social reasoning, since correct voting can be guesswork and effective lying may hide internal beliefs. MafiaScope provides an open testbed that converts Mafia into a measurement instrument for machine Theory of Mind: after each public utterance, every agent answers structured private probe questions scored against engine ground truth. An interactive visualizer tracks belief trajectories, supports impersonation and calibration panels, and enables counterfactual replay via forked steps. A 32-game DeepSeek study with 13,815 probe answers finds poor confidence calibration and includes a 30-fork replay workflow end-to-end, with released engine, viewer, and 200+ cross-model games.","MafiaScope: Non-Invasive, Time-Resolved Belief Probing for LLM Agents in Social Deduction Games*  \nIlia Karpov  \nHSE University  \n[karpovilia@gmail.com](karpovilia@gmail.com)  \narXiv :2607 . 10645v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nAn LLM agent’s public behaviour reveals little about its social reasoning: an agent that votes correctly may be guessing, and an agent that lies well leaves no trace of what it actually believes. We present MafiaScope, an open testbed that turns the social deduction game Mafia into a measurement instrument for machine Theory of Mind. After every public utterance, every agent privately answers a configurable set of structured probe questions;  \nthe answers never re-enter the game and are scored automatically against the ground truth the engine knows. An interactive visualizer renders the belief trajectories: impersonate mode shows the game as one agent sees it, panels chart timeline-aligned accuracy and calibration, and counterfactual replay forks any recorded step. In a 32-game DeepSeek case study with  \n13,815 parsed probe answers, stated confidence is poorly calibrated, with expected calibration error 0.17, agents over-predict being suspected  \n1.5 times, and a 30-fork replay experiment walks the counterfactual replay workflow end to end. Engine, viewer and a corpus of 200+ cross-model games are released under an open licence; live demo: [https://karpovilia](https://karpovilia). [github.io/mafiascope/](github.io/mafiascope/) ; screencast: [https:](https:)//[vimeo.com/1208920221](vimeo.com/1208920221) .  \n1 Introduction  \nMulti-agent LLM systems increasingly operate where success depends on modelling other minds: negotiation, collaboration, multi-party dialogue (Tan et al., 2023 ; Wei et al., 2023) . Static Theoryof-Mind (ToM) benchmarks probe such abilities with self-contained vignettes or scripted multiparty conversations (Le et al., 2019 ; Kim et al., 2023), but cannot show how an agent’s model of other agents evolves inside an adversarial interaction it partially causes (Ma et al., 2023 ; Riemer et al.,  \n* Paper has been submitted to the EMNLP 2026 System Demonstrations track.  \nFigure 1: The viewer on an English demo game: groundtruth graph and log in the centre, side panels are each agent’s private beliefs; node colour = assigned role, edge = suspicion. The ring on a living Mafioso,“hid N %”, is the share of the informed crowd missing it. Okabe-Ito palette with pattern redundancy.  \n2025) . Conversely, work that places LLMs inside social deduction games (Xu et al., 2023 ; O’Gara, 2023) typically evaluates only outcomes such as win rates and vote accuracy, leaving the agents’ beliefs a black box, or draws out reasoning inside the game context, where the questioning itself changes subsequent behaviour. Mafia 1 distils this challenge: an informed minority, the Mafia, who know eachother, covertly kills one player each night; the uninformed majority must unmask them through public discussion and daytime votes. Success on both sides hinges on modelling, and for the Mafia also managing, what the others believe.  \nWe demonstrate MafiaScope, a testbed built around one central idea, non-invasive, timeresolved belief probing: after every public utterance, privately question every agent with structured questions whose answers never re-enter the game,  \n1 [https://en.wikipedia.org/wiki/Mafia_](https://en.wikipedia.org/wiki/Mafia_) (party_game)  \nand score every answer against the ground truth the engine already knows.  \nMeasurement is therefore dense in time: every agent is probed after every public message, hundreds of belief snapshots per game, locating belief revision at the level of a single utterance. It is noninvasive in a checkable sense: probe answers never enter any agent’s persistent context; §4 spells out this guarantee and its limits. And it is self-scoring: Mafia’s hidden roles are known to the engine, so first-order beliefs (“Casey is Mafia”) are directly gradeable, and second-order beli","cbCaiid3VUhpxlO2","https://ap.wps.com/l/cbCaiid3VUhpxlO2","pdf",1654250,4,1,11,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What were key findings from the DeepSeek case study?\",\"answer\":\"Across 32 games and 13,815 parsed probe answers, stated confidence was poorly calibrated (ECE 0.17), with agents over-predict being suspected by about 1.5 times, and the counterfactual replay workflow was demonstrated end-to-end with 30 forks.\"}]",1784207639,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"mafiascope-non-invasive-time-resolved-belief-probing-for-llm-agents-in-social-deduction-games","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/mafiascope-non-invasive-time-resolved-belief-probing-for-llm-agents-in-social-deduction-games/85992/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What were key findings from the DeepSeek case study?","Question",{"text":75,"@type":76},"Across 32 games and 13,815 parsed probe answers, stated confidence was poorly calibrated (ECE 0.17), with agents over-predict being suspected by about 1.5 times, and the counterfactual replay workflow was demonstrated end-to-end with 30 forks.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]