[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85635-en":3,"doc-seo-85635-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85635,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules","Strategic voting is a key failure mode in collective choice, where voters can report ballots that differ from their true preferences to achieve more desirable outcomes. This paper presents an oracle-based benchmark that tests whether large language model (LLM) voters can discover and execute profitable manipulations. Each benchmark instance supplies true preference rankings, other voters’ ballots, a deterministic voting rule, and prompt framing conditions. An exact oracle enumerates all feasible reports, computes sincere and alternative outcomes, identifies profitable and optimal manipulations, and records the best achievable results. The benchmark covers multiple voting rules and prompt framings, using a fixed electorate size and balanced election instances to provide ground-truth calibration without subjective human grading.","Do Large Language Model Voters Strategize? An Oracle-Based Benchmark for Manipulation under Voting Rules  \narXiv :2606 .2 100 1v2 [ cs .GT] 10 Jul 2026  \nSeyed Pouyan Mousavi Davoudi  \nIndependent Researcher in AI and Statistics Tehran, Iran [spouyan.mousavi@gmail.com](spouyan.mousavi@gmail.com)  \nAmin Gholami Davodi  \nIndependent Researcher in AI and Statistics Tehran, Iran [a.g.davodi@gmail.com](a.g.davodi@gmail.com)  \nArshia  \nAlireza Amiri-Margavi  \nModel Risk Manager The Bank of New York New York, NY, USA [alireza.amirimargavi@Bny.com](alireza.amirimargavi@Bny.com)  \nHamidreza Hasani Balyani  \nAI Evaluation Engineer,  \nAmazon Lab126, Hardware Technology Organization Sunnyvale, CA, USA  \n[rezahsni@amazon.com](rezahsni@amazon.com)[ ](rezahsni@amazon.com)Gharagozlou*  \nMathematics & Statistics Department  \nUniversity of Minnesota Duluth  \nDuluth, MN, USA  \n[ghara027@d.umn.edu](ghara027@d.umn.edu)  \n*  \nEqual contribution.  \nAbstract  \nStrategic voting is a canonical failure mode for collective choice: a voter may obtain amore preferred outcome by reporting a ballot that differs from its true preferences. This paper introduces an oracle-based benchmark for testing whether large language model (LLM) voters can discover and execute such manipulations. Each instance gives an LLM voter a true preference ranking, the other voters’ ballots, a deterministic voting rule, and a prompt condition. An exact oracle enumerates every feasible report by the LLM voter, computes the sincere outcome, identifies all profitable reports, and records the best achievable outcome. The benchmark therefore supplies ground truth for strategic success without human labels or subjective grading of explanations. The benchmark covers plurality, Borda, approval, instant-runoff voting, and Copeland-style pairwise majority voting; prompt conditions separate sincere, strategic, civic, and expert framings. To keep the primary study defensible while preserving the main comparisons, the registered core design fixes a single electorate size, uses 600 balanced election instances, and produces 9,600 model–prompt responses when run with four model configurations and four prompt conditions. Because existing peer-reviewed work does not report manipulation discovery, optimal manipulation, false manipulation, near-miss, or invalid-ballot rates for this exact task, we do not impute LLM performance from unrelated studies. Instead, we report exact oracle-calibration baselines that bound and contextualize subsequent model results. By reducing strategic-voting behavior to exact counterfactual evaluation, the benchmark turns the question“Do LLM voters vote sincerely or strategically?” into a reproducible social-choice experiment.  \n1 Introduction  \nVoting rules are increasingly relevant to AI systems. LLM agents may be asked to participate in collective decisions, simulate voters, advise human voters, aggregate preferences from multiple models, or choose among candidate outputs generated by other systems. In these settings, a model’s  \nballot is not merely a text completion; it can change a collective outcome. Strategic voting is therefore a useful diagnostic for LLM agency: when the model is given preferences, a voting rule, and enough information to reason about the election, does it report sincerely or exploit a profitable misreport?  \nClassical social choice gives a sharp theoretical reason to expect strategic opportunities. The Gibbard–Satterthwaite theorem shows that every deterministic, onto, non-dictatorial social choice rule with at least three alternatives is manipulable at some profile [11, 14] . Computational social choice then asks whether finding such a manipulation is easy or hard under particular rules and candidate regimes [4–6, 8] . These literatures establish that manipulation exists and that its computational difficulty varies. They do not answer the empirical question raised by LLM agents: when an election profile is presented in natural language, can a mod","cbCaisLkksJe5BQz","https://ap.wps.com/l/cbCaisLkksJe5BQz","pdf",368873,3,1,19,"English","en",105,"# Introduction\n## Benchmark design and oracle calibration\n## Research questions\n## Contributions","[{\"question\":\"What does the oracle-based benchmark measure about LLM voters?\",\"answer\":\"It measures whether an LLM voter can discover and carry out profitable and optimal ballot manipulations relative to a deterministic voting rule, using exact oracle enumeration and ground-truth outcomes.\"},{\"question\":\"Which voting rules are included in the benchmark?\",\"answer\":\"The benchmark covers plurality, Borda, approval, instant-runoff voting, and Copeland-style pairwise majority voting.\"},{\"question\":\"How do prompt framings affect evaluation in the benchmark?\",\"answer\":\"Prompt conditions separate sincere, strategic, civic, and expert framings, allowing comparisons of how framing changes model behavior under the same voting tasks.\"}]",1784205162,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"do-large-language-model-voters-strategize-an-oracle-based-benchmark-for-manipulation-under-voting-rules","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/do-large-language-model-voters-strategize-an-oracle-based-benchmark-for-manipulation-under-voting-rules/85635/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the oracle-based benchmark measure about LLM voters?","Question",{"text":75,"@type":76},"It measures whether an LLM voter can discover and carry out profitable and optimal ballot manipulations relative to a deterministic voting rule, using exact oracle enumeration and ground-truth outcomes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which voting rules are included in the benchmark?",{"text":80,"@type":76},"The benchmark covers plurality, Borda, approval, instant-runoff voting, and Copeland-style pairwise majority voting.",{"name":82,"@type":73,"acceptedAnswer":83},"How do prompt framings affect evaluation in the benchmark?",{"text":84,"@type":76},"Prompt conditions separate sincere, strategic, civic, and expert framings, allowing comparisons of how framing changes model behavior under the same voting tasks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]