[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86028-en":3,"doc-seo-86028-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86028,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Opti Agent Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems","LLM agents increasingly target optimization tasks, yet existing benchmarks often assess them on pre-structured mathematical forms, skipping the hardest step: converting complex business requirements into correct models and solving them efficiently. Opti-Agent-Bench provides an end-to-end evaluation across the full optimization R&D pipeline—business-language understanding, mathematical modeling, algorithm selection and code implementation, and solution report generation. Three design pillars—anti-template traps, cross-module consistency checks, and an ORAC bi-level validity framework—reveal failures such as constraint omission, model-code inconsistency, and report-implementation divergence.","arXiv :2607 . 10768v 1 [ cs .AI] 12 Jul 2026  \nOpti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents  \non Real-World Business Problems  \nYongchang Fu 1 , Xinjie Huang 1 ,2 , Chengjun Dai 1 , Chengzhe Feng 1 , Junshao Zhang 1 , Hong Zhu 1  \n1 Ding Talk, Alibaba Group, Hangzhou, China  \n2 Zhejiang University, Hangzhou, China  \n{fuyongchang.fyc, chengjun.dcj, fengchengzhe.fcz, jiushao.zhangjs, [yisu}@alibaba-inc.com](yisu}@alibaba-inc.com)  \n[xinjiehuang@zju.edu.cn](xinjiehuang@zju.edu.cn)  \nAbstract  \nLLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathematical formulations that bypass the most critical challenge: translating complex business requirements into correct models and solve efficiently. We introduce Opti-Agent-Bench, an end-to-end benchmark that evaluates Large Language Models (LLMs) across the complete optimization R&D pipeline, from understanding business-language descriptions through mathematical modeling, algorithm selection, and code implementation, to solution report generation. Our design rests on three pillars: (1) businesssemantic authenticity with anti-template traps that defeat pattern matching; (2) modular evaluation with cross-module consistency checking across Problem Understanding, Formal Modeling, Implementation, and Reporting; and (3) the ORAC bi-level validity framework that simultaneously ensures task quality and scoring integrity. Across several industrialscale tasks spanning integer programming, robust optimization, stochastic programming, and non-convex optimization, we expose critical failure modes of current models, including constraint omission, model-code inconsistency, and report-implementation divergence, that remain invisible under conventional single-metric evaluation.  \n1 Introduction  \nOperations research and mathematical optimization underpin decision-making across industries—from supply chain logistics and production scheduling to financial portfolio design and energy network management. In practice, however, the path from a business stakeholder’s requirements to a deployed optimization solution is far from straightforward. It demands a complete R&D pipeline:  \nunderstanding the business context, identifying the underlying optimization structure, formulating a rigorous mathematical model, selecting and implementing an appropriate solution algorithm, and communicating the results through a clear, faithful report (see Figure 1) . Each link in this chain requires deep domain knowledge, mathematical sophistication, and engineering discipline.  \nThe rapid advancement of LLM-based agents has generated considerable excitement about automating this pipeline. Systems such as OptiMUS [AhmadiTeshnizi et al. , 2023] have demonstrated that LLM agents can, in favorable conditions, translate natural-language problem descriptions into solver code and iteratively debug their solutions. Meanwhile, foundation model approaches for mixed-integer linear programming (MILP) [Li et al. , 2025a] and knowledge-augmented formulation  \nFigure 1: A Pipeline for Solving Optimization Problems in Real-World Business Scenarios.  \nframeworks [Peng et al. , 2025] suggest that integrating LLMs with optimization tools can yield practically useful results. Yet a fundamental question remains largely unanswered: How reliably can these agents handle the full complexity of real-world optimization R&D?  \n1.1 The Evaluation Gap  \nCurrent benchmarks for optimization agents suffer from three interrelated limitations that systematically overestimate agent capabilities.  \nOver-Structured Problem Descriptions. Most existing benchmarks—including NLP4LP [AhmadiTeshnizi et al. , 2023], OptiBench [Yang et al. , 2025], and the NL4Opt competition [Ramamonjison et al. , 2023]—present problems that are already substantially structured. Task descriptions often use academic terminology (“minimize the total weighted tardiness,”“subject to capacity con","cbCaicBr58bCzdUq","https://ap.wps.com/l/cbCaicBr58bCzdUq","pdf",2305779,2,1,57,"English","en",105,"# Introduction\n## The Evaluation Gap\n# Opti-Agent-Bench Design\n## Business-Semantic Authenticity with Anti-Template Traps\n## Modular Evaluation with Cross-Module Consistency Checking\n## ORAC Bi-level Validity Framework\n# Benchmark Tasks and Findings\n## Failure Modes Revealed on Industrial-Scale Problems","[{\"question\":\"What gap in existing optimization-agent benchmarks does Opti-Agent-Bench target?\",\"answer\":\"It addresses the fact that most benchmarks rely on pre-structured mathematical formulations, avoiding the core challenge of turning real business requirements into correct models and solving them efficiently.\"},{\"question\":\"How does Opti-Agent-Bench evaluate the entire optimization R\\u0026D pipeline?\",\"answer\":\"It assesses LLMs end to end, from understanding business-language descriptions to formal mathematical modeling, algorithm selection and code implementation, and faithful solution report generation.\"},{\"question\":\"What kinds of model failures does Opti-Agent-Bench expose that single-metric evaluations can miss?\",\"answer\":\"It reveals critical failure modes including constraint omission, model-code inconsistency, and divergence between the generated report and the implemented solution.\"}]",1784207904,144,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"opti-agent-bench-benchmarking-end-to-end-optimization-rd-agents-on-real-world-business-problems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/opti-agent-bench-benchmarking-end-to-end-optimization-rd-agents-on-real-world-business-problems/86028/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What gap in existing optimization-agent benchmarks does Opti-Agent-Bench target?","Question",{"text":75,"@type":76},"It addresses the fact that most benchmarks rely on pre-structured mathematical formulations, avoiding the core challenge of turning real business requirements into correct models and solving them efficiently.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Opti-Agent-Bench evaluate the entire optimization R&D pipeline?",{"text":80,"@type":76},"It assesses LLMs end to end, from understanding business-language descriptions to formal mathematical modeling, algorithm selection and code implementation, and faithful solution report generation.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of model failures does Opti-Agent-Bench expose that single-metric evaluations can miss?",{"text":84,"@type":76},"It reveals critical failure modes including constraint omission, model-code inconsistency, and divergence between the generated report and the implemented solution.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]