[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81764-en":3,"doc-seo-81764-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81764,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","PHREEQC-MCQ-200 Tool-Augmented Scientific Simulator Agents Diagnostic Benchmark","Large language model agents increasingly connect to scientific software, but it is unclear when tool access improves reliability rather than simply adding complexity. PHREEQC-MCQ-200 is introduced as a benchmark for tool-augmented agents on deterministic aqueous geochemistry simulations, providing 200 multiple-choice questions from 21 validated PHREEQC scenarios. Agents must build simulator inputs, run PHREEQC, inspect structured outputs, and select final answers. Results show substantial accuracy gains, yet non-monotonic regressions and protocol-dependent output access effects.","PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents  \nKe Zhang  \nUniversity of California, Riverside  \n[kzhan153@ucr.edu](kzhan153@ucr.edu)  \nSahchit Chundur  \nUniversity of California, Irvine  \n[smchundu@uci.edu](smchundu@uci.edu)  \narXiv :2607 .00436v 1 [ cs .AI] 1 Jul 2026  \nMohammad Javad Qomi∗  \nUniversity of California, Irvine [mjaq@uci.edu](mjaq@uci.edu)  \nMaziar Raissi∗  \nUniversity of California, Riverside [maziar.raissi1@ucr.edu](maziar.raissi1@ucr.edu)  \nAbstract  \nLarge language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific computation more reliable rather than merely more complex. We introduce PHREEQC-MCQ-200, a benchmark for evaluating tool-augmented agents on deterministic aqueousgeochemistry simulations. The benchmark contains 200 multiple-choice questions derived from 21 validated PHREEQC scenarios, requiring agents to construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to final answers.  \nAcross multiple frontier and mid-tier model families, simulator access substantially improves aggregate accuracy, confirming that grounded execution is necessary for many scientific-computation tasks. However, the gains are not monotonic: tool-augmented agents also lose items they answered correctly without tools, revealing regressions that average accuracy alone hides. We further show that output-access protocol matters. A table-of-contents interface can reduce token cost while preserving or improving accuracy for stronger models, but it degrades performance for mid-tier models that cannot reliably navigate structured simulator outputs.  \nPHREEQC-MCQ-200 therefore frames scientific tool use as an end-to-end diagnostic problem rather than a simple tool-calling capability. We argue that evaluations of scientific agents should report not only accuracy, but also item-level retention, output-access sensitivity, trajectory failures, and where the computation chain breaks.  \n1 Introduction  \nBuilding tool-augmented LLM agents has become routine; evaluating whether tools actually make them more reliable on real scientific work has not. LLM agents are increasingly connected to external tools, including calculators, databases, code interpreters, and scientific software. Frameworks such as ReAct [Yao et al., 2023] and Toolformer [Schick et al., 2023] helped establish tool use asa core language-model capability, and scientific systems such as ChemCrow [Bran et al., 2024] and autonomous chemistry agents [Boiko et al., 2023] show that tools can extend model behavior beyond text-only prediction. For scientific computation, however, the central evaluation question isnot simply whether a model can call a tool. It is when tool augmentation improves a model’s ability  \n∗Corresponding authors.  \nPreprint.  \nto operate a deterministic scientific workflow, and when the tool loop itself becomes a source of errors [Zhang et al., 2026] . We study this on a deterministic scientific simulator and find a crossvendor capability-tier × tool-use interaction.  \nDeterministic scientific simulators make this question testable: simulator tasks provide executable evidence and reproducible ground truth. We study this through PHREEQC, an open-source simulator for aqueous geochemistry widely used by researchers and environmental consultants [Parkhurst and Appelo, 2013] . PHREEQC supports equilibrium speciation, mineral saturation, surface complexation, kinetic reactions, and one-dimensional reactive transport [Parkhurst and Appelo, 2013] . Its application scope spans groundwater and pollution studies [Appelo and Postma, 2004], geothermal-well scaling and corrosion [Bozau et al., 2015], and reactive-transport workflows [Parkhurst and Appelo, 2013] . The geothermal setting makes the practical stakes concrete: market estimates project global geothermal revenue approaching $10 billion by 2030 [Grand View Research, 2024], while","cbCaigcXzk3LAlFn","https://ap.wps.com/l/cbCaigcXzk3LAlFn","pdf",4022748,1,30,"English","en",105,"# Introduction\n## PHREEQC background and why determinism matters\n## PHREEQC-MCQ-200 benchmark design\n## Evaluation questions and diagnostics","[{\"question\":\"What problem does PHREEQC-MCQ-200 target for tool-augmented scientific agents?\",\"answer\":\"It evaluates whether adding tool access makes scientific computation more reliable in a deterministic workflow, not just more complex. The benchmark is designed to test end-to-end performance on simulator-driven tasks.\"},{\"question\":\"How is the benchmark constructed and how do agents answer questions?\",\"answer\":\"PHREEQC-MCQ-200 contains 200 multiple-choice items derived from 21 validated PHREEQC scenarios. For each item, agents must construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to a final answer.\"},{\"question\":\"Why are improvements from tool access not monotonic in the benchmark?\",\"answer\":\"Tool-augmented agents gain some items while losing others that they could answer correctly without tools. The paper argues that average accuracy alone can hide these regressions, so item-level retention and loss must be reported.\"}]",1784175982,76,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"phreeqc-mcq-200-tool-augmented-scientific-simulator-agents-diagnostic-benchmark","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/phreeqc-mcq-200-tool-augmented-scientific-simulator-agents-diagnostic-benchmark/81764/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does PHREEQC-MCQ-200 target for tool-augmented scientific agents?","Question",{"text":75,"@type":76},"It evaluates whether adding tool access makes scientific computation more reliable in a deterministic workflow, not just more complex. The benchmark is designed to test end-to-end performance on simulator-driven tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the benchmark constructed and how do agents answer questions?",{"text":80,"@type":76},"PHREEQC-MCQ-200 contains 200 multiple-choice items derived from 21 validated PHREEQC scenarios. For each item, agents must construct simulator inputs, execute PHREEQC, inspect structured outputs, and commit to a final answer.",{"name":82,"@type":73,"acceptedAnswer":83},"Why are improvements from tool access not monotonic in the benchmark?",{"text":84,"@type":76},"Tool-augmented agents gain some items while losing others that they could answer correctly without tools. The paper argues that average accuracy alone can hide these regressions, so item-level retention and loss must be reported.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]