[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83295-en":3,"doc-seo-83295-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83295,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier","From Solvers to Research argues that current AI for Mathematics systems—especially LLM-driven theorem provers in interactive theorem proving (ITP) environments—excel at proving well-defined problems but fall short for frontier research tasks. Open questions are often under-specified, multi-layered, and require novel, rigorous formal reasoning. The position paper surveys AI4Math methods, including datasets, auto-formalization, and proof synthesis, then analyzes limitations across data, relational structure, exploration, tool ecosystems, and human–AI collaboration, proposing a strategic roadmap to shift systems from solvers to research agents.","arXiv :2607 .07779v 1 [ cs .CL] 8 Jul 2026  \nFrom Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier  \nEric Jiang1,∗, Xiao Liang1,∗, Yikai Zhang1, Yingjia Wan1, Mengting Li1, Haikang Deng1, Alexander K Taylor1, Justin Baker1, Rushil Raghavan1, Junyi Zhang1, Ying Nian Wu1, Andrea L. Bertozzi1, Kai-Wei Chang1, Raghu Meka1,  \nMatthew Sottile2, Nanyun Peng1, Amit Sahai1, Terence Tao1, Wei Wang1  \n1University of California, Los Angeles  \n2Lawrence Livermore National Laboratory  \nAbstract  \nRecent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable success in formal proof generation for well-defined mathematical problems through Interactive Theorem Proving (ITP) languages. However, current systems remain fundamentally limited in tackling frontier research mathematics, such as discovering new theorems or resolving open conjectures, which are often open-ended, under-specified, and involve multiple layers of abstraction. We argue that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning. In this position paper, we provide a systematic review of the field, covering datasets, auto-formalization, and proof synthesis. More importantly, we identify core limitations of existing systems in serving as mathematical research agents, examining issues across datasets, relational structure, mathematical exploration, tool ecosystem, and human-AI collaboration, outlining a strategic road-map for the future of AI4Math.  \n Collection of Resources  \n1 Introduction  \nAI for Mathematics (AI4Math) has long been a central and foundational area of machine intelligence, reflecting the long-standing ambition to endow machines with rigorous formal mathematical reasoning capabilities. Early work in this field focused on neural-symbolic methods [282] designed to integrate neural pattern recognition with the structured logic of Interactive Theorem Proving (ITP) systems. These approaches have achieved notable success in high-accuracy proof synthesis within formal environments [274, 18, 124], but often rely on fixed, manually designed heuristics, limiting their scalability and applicability across diverse mathematical domains [1] .  \nRecently, the emergence of Large Language Models (LLMs) has led to remarkable progress in informal mathematical reasoning. Models like DeepSeek-R1 [57] and the o-series [106] have achieved strong performance on numerous benchmarks [164, 97] . However, these LLM reasoners that generate informal reasoning in natural language are fundamentally limited by the lack of precise, machine-checkable semantics, making their outputs prone to hallucinations [103] and precluding autonomous verification, a prerequisite for tackling open-ended mathematical research.  \nTo bridge this gap, research has been geared towards LLM-driven formal mathematical reasoning systems. By leveraging ITPs such as Lean [50] for rigorous verification, systems including DeepSeek-Prover [55, 267] and Seed-Prover [34] have set new standards for formal proof generation in competition-level mathemat-  \n∗ Equal contribution.  \nFor M. Sottile, this work was performed under the auspices of the U.S. Department of Energy by Lawrence Livermore National Laboratory under Contract DE-AC52-07NA27344 .  \nics [305] . In parallel, hybrid approaches [293, 38] that combine LLMs with geometrical deduction engines have surpassed human gold-medal performance on International Mathematical Olympiad (IMO) geometry problems. These advances highlight LLMs’ potential in generating formal proofs across diverse mathematical domains.  \nDespite these strides, we argue that current AI4Math systems still largely operate as solvers, excelling at isolated, well-defined proof generation rather than as researchers capable of expanding the bou","cbCaijVbMWcH0Ou1","https://ap.wps.com/l/cbCaijVbMWcH0Ou1","pdf",1865666,3,1,47,"English","en",105,"# Introduction\n## From solver-oriented systems to research agents\n# Collection of Resources\n## Abstract and overview\n# Unified analysis and taxonomy\n## Datasets, auto-formalization, and proof synthesis","[{\"question\":\"为什么现有基于LLM的形式化数学系统在研究前沿任务上仍受限？\",\"answer\":\"因为它们主要擅长在定义良好的问题上生成形式化证明，而研究前沿常涉及开放、欠指定且多层抽象的问题；此外，许多系统缺乏可机检的精确语义，从而难以实现自主验证与真正的研究级探索。\"},{\"question\":\"文中提出的核心转变是什么？\",\"answer\":\"从预定义的“problem-solver”转向面向前沿数学的“research agents”，使系统能够以严格形式化推理来应对开放式研究挑战。\"},{\"question\":\"本文在工作内容上主要覆盖哪些方面？\",\"answer\":\"围绕AI4Math给出系统综述，涵盖数据集、自动形式化与证明合成，并识别现有系统作为数学研究代理的关键限制，随后提出面向未来的路线图与改进方向。\"}]",1784186544,118,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-solvers-to-research-large-language-model-driven-formal-mathematics-at-the-research-frontier","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/from-solvers-to-research-large-language-model-driven-formal-mathematics-at-the-research-frontier/83295/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么现有基于LLM的形式化数学系统在研究前沿任务上仍受限？","Question",{"text":75,"@type":76},"因为它们主要擅长在定义良好的问题上生成形式化证明，而研究前沿常涉及开放、欠指定且多层抽象的问题；此外，许多系统缺乏可机检的精确语义，从而难以实现自主验证与真正的研究级探索。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文中提出的核心转变是什么？",{"text":80,"@type":76},"从预定义的“problem-solver”转向面向前沿数学的“research agents”，使系统能够以严格形式化推理来应对开放式研究挑战。",{"name":82,"@type":73,"acceptedAnswer":83},"本文在工作内容上主要覆盖哪些方面？",{"text":84,"@type":76},"围绕AI4Math给出系统综述，涵盖数据集、自动形式化与证明合成，并识别现有系统作为数学研究代理的关键限制，随后提出面向未来的路线图与改进方向。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]