[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83682-en":3,"doc-seo-83682-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83682,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents in Delegated Bargaining Under Privacy, Consent, Evidence, and Institutional Pressure","Personal agents increasingly negotiate on users’ behalf, potentially achieving agreements while harming users through privacy leakage, consent violations, unsupported advocacy, over-concession, failed escalation, or weak auditability. SovereignNegotiation-Bench is a trace-level, multi-turn benchmark for delegated bargaining under private utilities, disclosure constraints, evidence requirements, and institutional asymmetry. It separates agent-observable state from evaluator-only labels and jointly evaluates agreement success with user utility, privacy, consent, evidence grounding, concession discipline, escalation, and auditability.","arXiv :2607 .028 14v 1 [ cs .MA] 2 Jul 2026  \nSOVEREIGNNEGOTIATION-BENCH: EVALUATING USER-OWNED PERSONAL AGENTS IN DELEGATED BARGAINING UNDER PRIVACY, CONSENT, EVIDENCE , AND INSTITUTIONAL PRESSURE  \nDylan Zongmin Liu  \nStanford University  \n[zongminl@stanford.edu](zongminl@stanford.edu)  \nABSTRACT  \nPersonal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, escalating support disputes, requesting refunds, changing subscriptions, and negotiating deadlines or reimbursements. Existing negotiation benchmarks emphasize agreement, surplus, or strategic competence, but a user-owned agent can reach an agreement while harming the user through privacy leakage, consent violation, unsupported advocacy, over-concession, failed escalation, or poor auditability. We introduce SovereignNegotiation-Bench, a trace-level multi-turn benchmark for delegated personal-agent negotiation under private utilities, disclosure constraints, evidence requirements, and institutional asymmetry. The benchmark separates agent-visible observable state from evaluator-only labels and evaluates agreement success jointly with user utility, privacy, consent, evidence grounding, concession discipline, escalation, and auditability. We report an artifact-backed validation over 240 scenarios, 4 model families, 14 baselines, 13,440 frozen-prompt live trajectories, 61,135 parsed action rows, and a blinded 3-annotator audit over 300 items. The strongest agreement-maximizing baseline achieves the highest agreement rate but low user utility and high privacy/consent risk; FullSovereign does not maximize agreement, but obtains the best sovereign negotiation score by preserving utility, minimizing leakage, grounding claims, and reducing unauthorized commitments. The results show that agreement success is insufficient for user-owned negotiation agents.  \n1 INTRODUCTION  \nA personal agent that negotiates for a user is not merely a chatbot that bargains. It is a representative. It may know the user’s reservation value, private context, constraints, evidence, tolerance for escalation, and consent boundaries. This creates a failure mode absent from ordinary negotiation benchmarks:  \nthe agent can produce a deal and still violate the user’s interests. It might disclose a private medical reason to get a scheduling change, reveal a reservation price to accelerate a sale, accept a lowball settlement to end a support dispute, make unsupported claims in a refund appeal, or waive rights without authorization.  \nWe call this target sovereign negotiation: delegated bargaining that preserves the user’s current utility, privacy, consent, evidence standards, escalation rights, and auditability. The central thesis is that agreement rate is an incomplete metric for user-owned personal agents. Agreement-only evaluation can reward agents that are efficient for counterparties but poor representatives for users. SovereignNegotiation-Bench evaluates this target in two settings. Symmetric scenarios involve one user-owned personal agent negotiating with another personal agent, such as splitting shared costs or coordinating schedules. Asymmetric scenarios involve a user-owned personal agent facing a company or institution-side agent, such as a support escalation, refund request, platform appeal, subscription cancellation, or reimbursement claim. All policies see only observable negotiation state; private utilities, forbidden disclosures, acceptable evidence, concession thresholds, escalation labels, and internal  \npressure types are evaluator-only. The v3 prompt construction exposes only neutral counterparty profiles such as personal   counterparty and company   support   agent; strategy labels are not agent-visible.  \nContributions. (1) We define sovereign negotiation as an evaluation target for user-owned personal agents. (2) We introduce a no-oracle benchmark schema that separates observable negotiation state from hidden pri","cbCaiopAVgWI4BXB","https://ap.wps.com/l/cbCaiopAVgWI4BXB","pdf",201189,5,1,7,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Methodology\n# Experiments\n# Results\n# Contributions\n# Benchmark Design","[{\"question\":\"What problem does SovereignNegotiation-Bench target for user-owned personal agents?\",\"answer\":\"It targets the failure mode where an agent can reach an agreement while violating the user’s interests, such as leaking private information, breaking consent, making unsupported claims, or committing without authorization.\"},{\"question\":\"How does the benchmark evaluate beyond agreement rate?\",\"answer\":\"It evaluates agreement success jointly with user utility, privacy risk, consent compliance, evidence grounding, concession discipline, escalation behavior, and auditability, separating observable negotiation state from evaluator-only labels.\"},{\"question\":\"What settings does the benchmark cover?\",\"answer\":\"It includes symmetric scenarios (user-owned agent negotiating with another personal agent, e.g., shared costs or schedules) and asymmetric scenarios (user-owned agent facing a company or institution-side agent, e.g., support escalation, refunds, platform appeals, subscription cancellation, or reimbursement claims).\"}]",1784189710,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"sovereignnegotiation-bench-evaluating-user-owned-personal-agents-in-delegated-bargaining-under-privacy-consent-evidence-and-institutional-pressure","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/sovereignnegotiation-bench-evaluating-user-owned-personal-agents-in-delegated-bargaining-under-privacy-consent-evidence-and-institutional-pressure/83682/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does SovereignNegotiation-Bench target for user-owned personal agents?","Question",{"text":76,"@type":77},"It targets the failure mode where an agent can reach an agreement while violating the user’s interests, such as leaking private information, breaking consent, making unsupported claims, or committing without authorization.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the benchmark evaluate beyond agreement rate?",{"text":81,"@type":77},"It evaluates agreement success jointly with user utility, privacy risk, consent compliance, evidence grounding, concession discipline, escalation behavior, and auditability, separating observable negotiation state from evaluator-only labels.",{"name":83,"@type":74,"acceptedAnswer":84},"What settings does the benchmark cover?",{"text":85,"@type":77},"It includes symmetric scenarios (user-owned agent negotiating with another personal agent, e.g., shared costs or schedules) and asymmetric scenarios (user-owned agent facing a company or institution-side agent, e.g., support escalation, refunds, platform appeals, subscription cancellation, or reimbursement claims).","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]