[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82359-en":3,"doc-seo-82359-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82359,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","ProofCouncil: An LLM Agent for Solving Open Mathematical Problems","Large language models (LLMs) show growing capability for solving open mathematical problems, yet performance can be improved by using agentic workflows aligned with real mathematical practice. ProofCouncil is introduced as a mathematical agent built on an author–critic architecture, where an author iteratively writes and revises proofs while a critic evaluates and provides feedback. In the FirstProof challenge, it achieved top team performance, and on 30 researcher-supplied open problems it produced multiple fully correct solutions and useful partial progress. The paper also releases the agent-building library as open source.","arXiv :2607 .09474v 1 [ cs .AI] 10 Jul 2026  \nProofCouncil: An LLM Agent for Solving Open Mathematical Problems  \nJohannes Schmitt 1 , Tim Gehrunger 1 , Jasper Dekoninck 1 , Gergely B´erczi2 , Uri Kreitner 1 , Liam Price3 , David Holmes4  \n1 ETH Zurich, 2 Aarhus University, 3 Independent Researcher, 4 Leiden University Correspondence: johannes .schmitt@math .ethz .ch  \n [https://github.com/eth-sri/proof-council](https://github.com/eth-sri/proof-council)  \nAbstract  \nLarge language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be further improved through agentic workflows tailored to real-world mathematical practice. To this end, we introduce ProofCouncil, a mathematical agent that is designed to tackle open problems using an author-critic architecture.  \nProofCouncil served as a submission to the second batch of FirstProof, a challenge consisting of 10 real-world mathematical problems that agents must solve autonomously. Its submissions for 6 of the 10 problems were judged by the referees to be correct up to at most minor revisions, showing the best performance among participating teams. We also evaluate ProofCouncil on  \n30 open problems collected from mathematical researchers. Among the 21 solutions that received human feedback, 5 were judged completely correct, 2 more were judged promising pending final verification, and a further  \n8 contained useful partial progress. In this short paper, we describe the development of ProofCouncil and the agent-building library used to create it, which we release as open source to the community.  \n1 Introduction  \nLarge language models (LLMs) have recently demonstrated remarkable capabilities in mathematics by solving open problems of varying interest (Alon et al. , 2026; Schmitt, 2025; Dixit et al. , 2026) . This progress has been accompanied by the development of advanced agents designed for mathematical problem-solving (Dixit et al. , 2026; Zhao et al. , 2026; Zheng et al. , 2026; Ju et al. , 2026), further highlighting the potential of LLMs in this domain. In this context, the second iteration of the FirstProof challenge (Abouzaid et al. , 2026a;b) was launched to evaluate AI systems on a fixed set of 10 problems with no publicly available solution. Participants were required to submit an agent that autonomously attempted all 10 problems within a 24-hour window, using only publicly available models and tools.  \nIn this work, we present ProofCouncil, a mathematical agent submitted to the FirstProof challenge. As shown in Fig. 1, ProofCouncil follows an author-critic architecture: an author agent iteratively writes and revises a proof, while a critic agent evaluates each version and provides feedback. At every round, the author may also request targeted assistance from two auxiliary sources: a council of additional LLMs outside the main author-critic loop, anda compute agent for computer algebra system (CAS) computations. Their responses are provided to the author in the next iteration. To maintain consistency and allow the critic to track progress, the critic generally retains its conversation history across rounds, but this history is reset every k rounds to provide a more independent review. If a freshly initialized critic accepts the proof, the proof is returned to the user.  \nWe open-source both ProofCouncil and the underlying agent-building library. The library represents agentic systems as conditional directed acyclic graphs (DAGs), enabling flexible structures of calls between different agents and models. It also provides a simple interface for  \nFigure 1: Overview of ProofCouncil. The author agent iteratively edits the proof in response to feedback from a critic. Every k rounds, the stateful critic is reset. The author may optionally request help from other LLMs or a compute node.  \nimplementing new agents and workflows, enabling future researchers to build on our work and create their own agentic systems.  \n","cbCaidMnHUdcZFiH","https://ap.wps.com/l/cbCaidMnHUdcZFiH","pdf",1318875,6,1,25,"English","en",105,"# Introduction\n# ProofCouncil\n## ProofCouncil Workflow","[{\"question\":\"What is ProofCouncil and how does it work for solving open mathematical problems?\",\"answer\":\"ProofCouncil is a mathematical agent that uses an author–critic architecture. The author agent iteratively drafts and revises a proof, while the critic agent evaluates each version and returns targeted feedback. Auxiliary assistance can be requested from other LLMs and a CAS compute agent in later iterations.\"},{\"question\":\"How did ProofCouncil perform in the FirstProof challenge?\",\"answer\":\"ProofCouncil was submitted to the second batch of FirstProof with 10 real-world mathematical problems. For 6 of the 10 problems, referees judged the submissions correct with at most minor revisions, which was the best performance among the participating teams.\"},{\"question\":\"What evaluation was done beyond the challenge, and what results were reported?\",\"answer\":\"The paper evaluates ProofCouncil on 30 open problems collected from mathematical researchers. Among 21 solutions that received human feedback, 5 were judged completely correct, 2 were promising pending final verification, and 8 provided useful partial progress.\"}]","ProofCouncil: An LLM Agent for Solving Open Mathematical Problems | PDF",1784179881,63,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"proofcouncil-an-llm-agent-for-solving-open-mathematical-problems","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/proofcouncil-an-llm-agent-for-solving-open-mathematical-problems/82359/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-30","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What is ProofCouncil and how does it work for solving open mathematical problems?","Question",{"text":77,"@type":78},"ProofCouncil is a mathematical agent that uses an author–critic architecture. The author agent iteratively drafts and revises a proof, while the critic agent evaluates each version and returns targeted feedback. Auxiliary assistance can be requested from other LLMs and a CAS compute agent in later iterations.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How did ProofCouncil perform in the FirstProof challenge?",{"text":82,"@type":78},"ProofCouncil was submitted to the second batch of FirstProof with 10 real-world mathematical problems. For 6 of the 10 problems, referees judged the submissions correct with at most minor revisions, which was the best performance among the participating teams.",{"name":84,"@type":75,"acceptedAnswer":85},"What evaluation was done beyond the challenge, and what results were reported?",{"text":86,"@type":78},"The paper evaluates ProofCouncil on 30 open problems collected from mathematical researchers. Among 21 solutions that received human feedback, 5 were judged completely correct, 2 were promising pending final verification, and 8 provided useful partial progress.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]