[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84020-en":3,"doc-seo-84020-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84020,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Strategic Bargaining in Multi-Buyer Markets Reinforcement Learning from Verifiable Rewards for LLM Negotiations","Strategic bargaining addresses agreements under private information in multi-buyer settings where a single seller negotiates concurrently with heterogeneous buyers who hold hidden budgets. Standard large language models can generate language but do not behave as effective economic decision-makers, often failing to explore the buyer pool and instead fixating on current highest bids. The work introduces a reinforcement learning from verifiable rewards (RLVR) training approach that anchors rewards to objective economic outcomes, yielding learned market discovery and surplus extraction strategies. Experiments show higher surplus and robust generalization to unseen buyer styles and budget distributions.","arXiv :2607 .05863v 1 [ cs .LG] 7 Jul 2026  \nStrategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for  \nLLM Negotiations  \nShuze Daniel Liu  \nInstitute for Data, Systems, and Society, Massachusetts Institute of Technology, [shuzel@mit.edu](shuzel@mit.edu); Mitch Daniels School of Business, Purdue University, [daniel.liu@purdue.edu](daniel.liu@purdue.edu)  \nClaire Chen  \nThe Division of Physics, Mathematics and Astronomy, California Institute of Technology, [clairechen@caltech.edu](clairechen@caltech.edu)  \nJiabao Sean Xiao  \nDepartment of Computing and Mathematical Sciences, California Institute of Technology, [seanxiao@caltech.edu](seanxiao@caltech.edu)  \nXin Chen  \nH. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology, [xin.chen@isye.gatech.edu](xin.chen@isye.gatech.edu)  \nDavid Simchi-Levi  \nInstitute for Data, Systems, and Society, Department of Civil and Environmental Engineering, Operations Research Center,  \nMassachusetts Institute of Technology, [dslevi@mit.edu](dslevi@mit.edu);  \nMitch Daniels School of Business, Purdue University, [dslevi@purdue.edu](dslevi@purdue.edu)  \nNegotiation is a fundamental strategic interaction in management science, characterized by agents attempting to reach agreements while protecting private information, such as reservation costs and hidden valuations. A prevalent yet complex scenario involves a single seller negotiating concurrently with multiple buyers, each possessing heterogeneous, private budgets. In such settings, constrained by a limited number of communication turns, the seller must balance exploring the broader market to discover the highest valuation with concentrating sufficient turns on a single target buyer to secure the best possible outcome. Our analysis reveals a significant gap in standard Large Language Models (LLMs): while these models are linguistically proficient, they fail to act as effective economic decision-makers. Specifically, they exhibit a failure to explore the buyer pool, often fixating on the current highest bid rather than strategically investigating the market to discover latent high valuations.  \nIn this paper, we propose a specialized training recipe using Reinforcement Learning from Verifiable Rewards (RLVR) . By anchoring the reward function to objective economic outcomes, the strategic balance between market discovery and surplus extraction emerges natively through the learning process. Our results demonstrate that the trained seller undergoes a multi-stage strategic evolution, learning to leverage price anchoring and strategic probing to identify more profitable counterparties. The agent extracts substantially higher surplus than frontier models by both improving its persuasive bargaining skills and consistently closing deals with high-value buyers. Finally, we show that our seller strategies generalize robustly to unseen buyer negotiation styles and budget distributions.  \nKey words: LLM agents; reinforcement learning with verifiable rewards; negotiation; multi-agent bargaining; surplus extraction  \n1 . Introduction  \nNegotiation is a foundational process in management science and operations research, serving asthe primary mechanism for value discovery and resource allocation in environments where prices are not static (Simchi-Levi et al. 2005 , Backus et al. 2020) . In fields ranging from procurement and supply chain contracting to dispute resolution and online marketplaces, agents must strategically interact to coordinate and realize gains from trade (Cachon and Lariviere 2001 , Desai and Purohit 2004) . Within management science literature, the analysis of these strategic interactions focuses heavily on aligning decentralized incentives and resolving information gaps between trading partners (Nagarajan and Soˇsi´c 2008 , Lovejoy 2010) .  \nA central challenge in any negotiation is that agents possess private information—specifically, a seller has a minimum cost the","cbCaijabXqHxQya8","https://ap.wps.com/l/cbCaijabXqHxQya8","pdf",1438995,6,1,53,"English","en",105,"# Introduction\n## Negotiation under private information\n## Concurrent multi-buyer negotiation and exploration–exploitation trade-off\n## Limitations of standard LLMs in economic decision-making\n## Proposed RLVR training approach and outcomes","[{\"question\":\"What problem does the paper focus on in multi-buyer negotiations?\",\"answer\":\"It studies how a single seller negotiates concurrently with multiple buyers who have heterogeneous, private budgets while limited communication turns force an exploration–exploitation trade-off.\"},{\"question\":\"What limitation of standard large language models is highlighted?\",\"answer\":\"Although LLMs are linguistically capable, they do not act as effective economic decision-makers and often fail to explore the buyer pool, fixating on current highest bids instead.\"},{\"question\":\"How does the proposed RLVR training recipe work conceptually?\",\"answer\":\"It uses reinforcement learning from verifiable rewards by anchoring the reward function to objective economic outcomes, so the agent learns strategies that balance market discovery and surplus extraction.\"}]",1784192051,134,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"strategic-bargaining-in-multi-buyer-markets-reinforcement-learning-from-verifiable-rewards-for-llm-negotiations","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/strategic-bargaining-in-multi-buyer-markets-reinforcement-learning-from-verifiable-rewards-for-llm-negotiations/84020/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper focus on in multi-buyer negotiations?","Question",{"text":76,"@type":77},"It studies how a single seller negotiates concurrently with multiple buyers who have heterogeneous, private budgets while limited communication turns force an exploration–exploitation trade-off.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What limitation of standard large language models is highlighted?",{"text":81,"@type":77},"Although LLMs are linguistically capable, they do not act as effective economic decision-makers and often fail to explore the buyer pool, fixating on current highest bids instead.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed RLVR training recipe work conceptually?",{"text":85,"@type":77},"It uses reinforcement learning from verifiable rewards by anchoring the reward function to objective economic outcomes, so the agent learns strategies that balance market discovery and surplus extraction.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]