[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83151-en":3,"doc-seo-83151-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83151,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics","Recent advances in AI for Mathematics have focused mainly on autoformalization and theorem proving, leaving Computer Algebra Systems (CAS) largely underexplored in agentic LLM workflows. This work proposes a ReAct-style agent that combines LLM reasoning with verifiable feedback from SageMath and uses Context7 for current documentation. Frontier models are evaluated on RealMath under a computational-math research loop. A benchmark refinement with multi-step postprocessing and multi-stage validation improves extraction quality. Experiments show sizable SageMath gains across models, narrowing the gap between open and closed models.","Evaluating SageMath-Augmented LLM Agents for Computational and  \nExperimental Mathematics  \nPavel Snopov 1 * German Magai 2 *  \narXiv :2607 .06820v 1 [ cs .AI ] 7 Jul 2026  \nAbstract  \nRecent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a computational-mathematics research loop. We also propose a refinement to the RealMath benchmark by introducing a multi-step postprocessing procedure and a multi-stage validation pipeline, both of which improve the quality and reliability of the extracted problem set. Our experiments reveal substantial performance gains from SageMath access across all evaluated models on +9.7 pp on average, the gains range from 1.5 pp to 27.8 pp and narrow the gap between open-weight and closed models. Qwen 3 .7-Max benefits from SageMath the most, while GPT-5.5 achieves the highest solve rate of 75 .2% and the lowest token usage among tool-enabled configurations. Our findings suggest that CAS-augmented agents represent a promising direction for assisting mathematicians in computational exploration, and we believe that this work is a step towards automated conjecture discovery. The project repository is available online. 1  \n*Equal contribution 1 School of Mathematical and Statistical Sciences, The University of Texas Rio Grande Valley, USA 2Noeon Research, Tokyo, Japan. Correspondence to: Pavel Snopov \u003C[paul.snopov@utrgv.edu](paul.snopov@utrgv.edu) >, German Magai \u003Cger[man@noeon.ai](man@noeon.ai) > .  \nPreprint. July 9, 2026.  \n1 [https://github.com/Snopoff/Evaluating](https://github.com/Snopoff/Evaluating)SageMath-Augmented-LLM-Agents-forComputational-and-Experimental-Mathematics  \n1. Introduction  \nRecent progress in LLMs and agentic systems that integrate LLM reasoning with deterministic, verifiable tool backends such as compilers, theorem provers, SAT/SMT solvers, type checkers, and physical simulators has established a new neuro-symbolic paradigm in which generative reasoning is partially grounded in verifiable feedback. By coupling LLM reasoning with components that provide ground-truth signals, this paradigm enables a class of tools capable of automating tasks that previously required substantial expert effort and of producing answers with a level of reliability that purely generative approaches cannot achieve.  \nA particularly active application of this paradigm is in mathematics. Substantial progress has been made on autoformalization, the translation of mathematical statements into the formal programs written in proof-assistant languages such as Lean, Coq, and Isabelle (Wu et al., 2022), as well as on automated theorem proving, where LLM-based systems guide the search for formal proofs (Lin et al., 2025), attain gold-medal-level performance on the 2025 International Mathematical Olympiad with formally verified solutions (Achim et al., 2025), and and have recently autonomously resolved several open problems from the Erds collection (Sothanaphan, 2026 ; Tsoukalas et al., 2026) including the well-known unit-distance conjecture (OpenAI, 2026) . These developments establish the integration of LLMs with formal proof assistants as a promising direction for mathematical reasoning, focused on proving and formalizing already specified statements.  \nIn many areas of mathematics, Computer Algebra Systems (CAS) and symbolic engines are routinely used for hypothesis exploration, candidate validation, and counterexample search. A common workflow in modern mathematical research, particularly in computational areas such as combinatorial commutative algebra, algeb","cbCaish7n30cao22","https://ap.wps.com/l/cbCaish7n30cao22","pdf",7472614,4,1,37,"English","en",105,"# Introduction\n## Neuro-symbolic agentic systems in mathematics\n## Role of CAS in computational research workflows\n## Motivation and research questions\n## Contributions and evaluation setup","[{\"question\":\"What motivates introducing SageMath into LLM agent workflows for mathematics?\",\"answer\":\"The document argues that prior work emphasized autoformalization and theorem proving, while CAS-based symbolic experimentation and executable verification were comparatively underexplored. SageMath enables verifiable feedback aligned with computational research workflows.\"},{\"question\":\"How is the proposed LLM agentic setup designed?\",\"answer\":\"It uses a ReAct-style agent that couples LLM reasoning with verifiable feedback from SageMath, and incorporates Context7 for up-to-date documentation. The evaluation emulates a mathematician’s computational loop.\"},{\"question\":\"What benchmark modifications improve problem-set quality and reliability?\",\"answer\":\"The work refines RealMath with a multi-step postprocessing procedure and a multi-stage validation pipeline. These steps improve the quality and reliability of the extracted problem set.\"}]",1784185622,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-sagemath-augmented-llm-agents-for-computational-and-experimental-mathematics","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/evaluating-sagemath-augmented-llm-agents-for-computational-and-experimental-mathematics/83151/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What motivates introducing SageMath into LLM agent workflows for mathematics?","Question",{"text":75,"@type":76},"The document argues that prior work emphasized autoformalization and theorem proving, while CAS-based symbolic experimentation and executable verification were comparatively underexplored. SageMath enables verifiable feedback aligned with computational research workflows.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the proposed LLM agentic setup designed?",{"text":80,"@type":76},"It uses a ReAct-style agent that couples LLM reasoning with verifiable feedback from SageMath, and incorporates Context7 for up-to-date documentation. The evaluation emulates a mathematician’s computational loop.",{"name":82,"@type":73,"acceptedAnswer":83},"What benchmark modifications improve problem-set quality and reliability?",{"text":84,"@type":76},"The work refines RealMath with a multi-step postprocessing procedure and a multi-stage validation pipeline. These steps improve the quality and reliability of the extracted problem set.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]