[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84985-en":3,"doc-seo-84985-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84985,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning","Mathematical reasoning benchmarks often assume full information or mix reasoning with tools and retrieval. MIRA-Math isolates a narrower diagnostic ability: solving problems with a unique answer when the solver’s view is missing exactly one necessary atomic fact. The solver must request the missing fact in natural language within a strict budget and incorporate the returned fact into an exact final answer. A constrained responder shares only the dataset fact when the request matches, otherwise declines; evaluation uses deterministic generation, validation, and verification over 2,310 instances.","arXiv :2607 .0739 1v 1 [ cs .AI] 8 Jul 2026  \nMIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning  \nCharbel Al Bateh [charbel.albateh@lau.edu](charbel.albateh@lau.edu)  \nDepartment of Electrical and Computer Engineering Lebanese American University  \nByblos, Lebanon  \nSamer Saab [Jr.](Jr. samer.saabjr@lau.edu.lb)[ samer.saabjr@lau.edu.lb](Jr. samer.saabjr@lau.edu.lb)  \nDepartment of Electrical and Computer Engineering Lebanese American University  \nByblos, Lebanon  \nAbstract  \nMathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and longhorizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathematical problems whose full latent state has a unique answer, but whose solver-facing view is missing exactly one necessary atomic fact. The solver must request the missing information in natural language under a strict budget and then integrate the returned fact into an exact final answer. A fixed constrained LLM responder sees only the dataset-provided atomic fact and must either offer the quoted fact when the request matches it, or decline otherwise. Thus, instance generation, typed hint specifications, validation, and final-answer verification are deterministic, while request metrics are measured under a fixed LLM-mediated responder channel. MIRA-Math contains 2,310 generated instances from 22 typed mathematical families spanning algebra, probability, linear systems, discrete structures, signal processing, Markov chains, circuits, interpolation, and numerical boundary-value problems. Experiments across frontier and small models show that request success and final-answer accuracy are separable: models may ask for the right fact yet fail the downstream computation, or fail before obtaining the canonical hint. We release generators, verifiers, prompts, run metadata, and dataset documentation to support reproducible evaluation of minimal information requesting in mathematical reasoning.  \nKeywords: benchmark, dataset generation, mathematical reasoning, information acquisition, clarification, partial observability, large language models  \n1 Introduction  \nLarge language models are increasingly evaluated on mathematical reasoning, tool use, and interactive problem solving. Yet many high-profile math benchmarks assess only the final answer under full information (Cobbe et al., 2021; Hendrycks et al., 2021; Balunovi´c et al. , 2025), while broad interactive benchmarks mix multiple sources of difficulty, including toolselection, retrieval, web navigation, external APIs, and long-context management (Mialonet al., 2023; Liu et al., 2023; Qin et al., 2023) . These settings are valuable, but they make it difficult to isolate a basic bottleneck that appears whenever a model does not initially possess all relevant information: can the model recognize the missing fact, ask for that fact precisely, and use the answer correctly?  \n©2026 Al Bateh and Saab.  \nAl Bateh and Saab  \nMIRA-Math is designed to isolate this capability. Each benchmark instance is generated from a complete mathematical state with a unique answer, but the solver model receives only a private view that is deliberately insufficient, missing exactly one necessary atomic fact. A fixed information-holder model receives only the missing fact and is constrained by a structured-output protocol: if the solver’s request semantically matches the information it holds, it returns that fact; otherwise, it returns a declination stating that it does not have the requested information. The information holder does not solve, explain, negotiate, or volunteer extra hints. The benchmark therefore measures minimal information requesting under a controlled responder channel, not open-ended multi-agent collaboration.  \nThis framing is intentionally data-centric. Following the emphasis on public benchmark infrastructu","cbCair26oCMWsCvq","https://ap.wps.com/l/cbCair26oCMWsCvq","pdf",1295839,2,1,62,"English","en",105,"# Introduction\n## Final-answer mathematical reasoning is not enough\n## MIRA-Math framing and protocol\n# Related Work","[{\"question\":\"What problem does MIRA-Math measure in mathematical reasoning?\",\"answer\":\"It measures whether a model can identify a missing atomic fact, request it precisely in natural language, and then use the received fact to produce the correct final answer.\"},{\"question\":\"How is the responder in MIRA-Math constrained?\",\"answer\":\"A fixed constrained LLM responder sees only the single atomic fact provided by the dataset and follows a structured offer/decline protocol based on whether the solver’s request matches.\"},{\"question\":\"What does MIRA-Math evaluate and why are the metrics separable?\",\"answer\":\"Experiments separate request success from final-answer accuracy, showing that models may ask for the right fact yet still fail the downstream computation, or fail before obtaining the canonical hint.\"}]",1784200056,156,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mira-math-a-benchmark-for-minimal-information-requesting-and-mathematical-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mira-math-a-benchmark-for-minimal-information-requesting-and-mathematical-reasoning/84985/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MIRA-Math measure in mathematical reasoning?","Question",{"text":75,"@type":76},"It measures whether a model can identify a missing atomic fact, request it precisely in natural language, and then use the received fact to produce the correct final answer.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the responder in MIRA-Math constrained?",{"text":80,"@type":76},"A fixed constrained LLM responder sees only the single atomic fact provided by the dataset and follows a structured offer/decline protocol based on whether the solver’s request matches.",{"name":82,"@type":73,"acceptedAnswer":83},"What does MIRA-Math evaluate and why are the metrics separable?",{"text":84,"@type":76},"Experiments separate request success from final-answer accuracy, showing that models may ask for the right fact yet still fail the downstream computation, or fail before obtaining the canonical hint.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]