[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86335-en":3,"doc-seo-86335-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86335,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","AdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification","Large language models demonstrate strong results on high-school and competition-level math, but advanced mathematical reasoning remains insufficiently measured. Existing benchmarks often use limited disciplinary coverage and rely on final-answer correctness or coarse grading, leaving proof validity and reasoning quality under-evaluated. AdvancedMathBench introduces ProverBench (245 UG and QE problems) with an expert-trained automatic verification pipeline, plus VerifierBench (888 trajectories) to assess proof validity judging and verification rationales.","arXiv :2607 . 1 1849v 1 [ cs .CL] 13 Jul 2026  \n2026-7-14  \nAdvancedMathBench: A Benchmark Suite for Advanced Mathematical Proof Generation and Verification  \nLingkai Kong1,2 , Zijian Wu1,3 , Yuzhe Gu1,2 , Haiteng Zhao1 , Wenyong Huang1,2 , Shuang Sun1,2 , Zhicheng Xiong1,2 , Xiaotian Zhang1,2 , Shuya Zhao1,2 , Yan Wang2 , Disheng Xu4 , Wenwei Zhang1† and Kai Chen1†  \n1 Shanghai AI Laboratory, 2 Shanghai Jiao Tong University, 3 MMLab, The Chinese University of Hong Kong, 4 Great Bay University  \nLarge language models (LLMs) have achieved remarkable performance on high-school and competition-level mathematics, yet their capabilities on advanced mathematics remain poorly understood. Existing benchmarks, however, fall short in both scope and evaluation granularity: they provide limited disciplinary coverage and often rely on final-answer correctness or coarse judgments, leaving the validity of the reasoning process inadequately assessed. To bridge this gap, we introduce AdvancedMathBench, a benchmark suite designed to evaluate the reasoning capabilities of LLMs on advanced mathematics. Its core proof-generation benchmark, ProverBench, contains 245 problems spanning undergraduate (UG) and doctoral qualifying-exam (QE) levels. To provide reliable evaluation of the proofs, we develop a dedicated automatic verification pipeline trained on large-scale expert annotations to produce both correctness verdicts and fine-grained assessments of proof errors, which exhibits strong agreement with human experts on held-out proof trajectories. We further introduce VerifierBench, consisting of 888 model-generated proof trajectories paired with expert ground truth, to evaluate whether models can correctly judge proof validity and provide sound verification rationales. Experiments show that AdvancedMathBench remains challenging for frontier models. On proof generation, the best-performing model, GPT-5.5-xhigh, achieves only 64.5 and 48.9 on the UG and QE splits, respectively, indicating substantial room for improvement on advanced mathematical proof construction. On proof verification, the best model only attains a Balanced F1 of only 65.1 suggesting that critical error detection remains a major bottleneck in applying LLMs for proof verification.  \n1. Introduction  \nLarge language models (LLMs) have achieved remarkable progress on mathematical reasoning, especially on high-school and competition-level benchmarks [2, 6, 10, 14, 15, 44] . Yet their capabilities on advanced mathematics remain much less understood. Because advanced mathematics requires the model construct trajectory with rigorous intermediate claims, rather than captured by the final short-answer prediction. Mathematical proof therefore provides a natural stress test for evaluating whether models can reason beyond producing the correct final answer.  \nExisting benchmarks provide limited support for evaluating this capability. Most benchmarks still emphasize competition-style problems, and remain limited in coverage for undergraduate, graduate, and research levels of mathematics [5, 8, 9, 12, 39] . More importantly, many benchmarks are evaluated primarily through final-answer checking or coarse solution matching, which cannot determine whether a proof is mathematically valid. Recent work has explored process verification and LLM-as-Judge [4, 20, 23, 25, 31, 40, 41], but them still suffer from bias and inconsistency. Therefore, current evaluations may lead to an inaccurate estimation of the mathematical capabilities of a model, as they overlook the validity of the underlying reasoning process.  \nThis paper aims to provide a more rigorous foundation for measuring model behavior in advanced mathemat-  \n† Corresponding authors  \nPerformance Score (%)  \n100  \n90  \n80  \n70  \n60  \n50  \n40  \n30  \n20  \n10  \nMeta-Verification TNR  \n100  \n90  \n80  \n70  \n60  \n50  \n40  \n30  \n65 70 75 80 85 90 95 100  \nMeta-Verification TPR  \nFigure 1: AdvancedMathBench exposes complementary weaknesses in advanced ma","cbCaihekOUoStX0x","https://ap.wps.com/l/cbCaihekOUoStX0x","pdf",1436919,5,1,20,"English","en",105,"# Introduction\n## Motivation for advanced proof evaluation\n## AdvancedMathBench design and components\n## Evaluation results and remaining challenges","[{\"question\":\"Why do existing math benchmarks fail to assess advanced mathematical proof capability?\",\"answer\":\"They often emphasize final-answer correctness or coarse solution matching, which cannot reliably determine whether the reasoning steps form a mathematically valid proof.\"},{\"question\":\"What are the two main components of AdvancedMathBench?\",\"answer\":\"ProverBench evaluates proof generation, while VerifierBench evaluates proof verification, including validity judgment and justification of verification.\"},{\"question\":\"How is proof quality evaluated for model-generated proofs?\",\"answer\":\"A dedicated automatic verification pipeline is trained on large-scale expert annotations to output correctness verdicts, fine-grained error assessments, and error localization, aligning well with human expert evaluation.\"}]",1784210542,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"advancedmathbench-a-benchmark-suite-for-advanced-mathematical-proof-generation-and-verification","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/advancedmathbench-a-benchmark-suite-for-advanced-mathematical-proof-generation-and-verification/86335/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do existing math benchmarks fail to assess advanced mathematical proof capability?","Question",{"text":76,"@type":77},"They often emphasize final-answer correctness or coarse solution matching, which cannot reliably determine whether the reasoning steps form a mathematically valid proof.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the two main components of AdvancedMathBench?",{"text":81,"@type":77},"ProverBench evaluates proof generation, while VerifierBench evaluates proof verification, including validity judgment and justification of verification.",{"name":83,"@type":74,"acceptedAnswer":84},"How is proof quality evaluated for model-generated proofs?",{"text":85,"@type":77},"A dedicated automatic verification pipeline is trained on large-scale expert annotations to output correctness verdicts, fine-grained error assessments, and error localization, aligning well with human expert evaluation.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":20,"slug":136},19,"General","general"]