[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85716-en":3,"doc-seo-85716-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85716,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language","PHITSBench introduces an execution-scored benchmark for AI-assisted input generation for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). The benchmark contains 282 transport-scorable tasks across three workflow categories: parameter editing, syntax repair, and full simulation generation from natural-language descriptions. Evaluation uses a Composite Metric Score combining execution success and agreement with reference transport observables. Experiments with GPT-5.4 settings show strong editing and repair performance, while scratch simulation generation initially fails. A machine-readable PHITS knowledge catalog and agentic execution improve reproduce-track success to 57% and up to 66–73%, respectively.","arXiv :2607 .09789v 1 [ cs .AI] 8 Jul 2026  \nARTICLE  \nPHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language  \nXianglin Ji and Svetlana V. Boriskina  \nDepartment of Mechanical Engineering, Massachusetts Institute of Technology, Cambridge, MA 02139, USA  \nARTICLE HISTORY  \nCompiled July 14, 2026  \nABSTRACT  \nWe introduce PHITSBench, an execution-scored benchmark for the Monte Carlo  \nParticle and Heavy Ion Transport code System (PHITS) . PHITSBench comprises  \n282 transport-scorable tasks spanning three common workflow categories: parameter editing (Edit ), syntax repair (Repair ), and complete simulation generation from natural-language descriptions (Reproduce) . Each task is evaluated using a Composite Metric Score that combines execution success with agreement between generated and reference transport observables. Using PHITSBench, we evaluate five GPT-5.4-based configurations ranging from zero-shot prompting to knowledge-augmented and agentic workflows. Without domain-specific knowledge, the model performs well on editing and repair tasks (95% and 70% success, respectively) but fails to generate correct simulations from scratch (0% success on the Reproduce track) . A structured, machine-readable PHITS knowledge catalog, supplied alongside the user manual, raises single-shot Reproduce-task success to 57% . Agentic execution provides a further improvement to 66-73%, but at increased computational cost. Failure analysis shows that the remaining errors are dominated by incorrect selection and configuration of physical observables rather than syntax generation. These results suggest that future progress in AI-assisted radiation-transport modeling will depend as much on machinereadable knowledge bases, curated domain-training datasets, and execution-grounded evaluation environments as on advances in foundation models themselves.  \nKEYWORDS  \nPHITS; Monte Carlo radiation transport; large language models; benchmark; input file generation; agentic systems  \n1. Introduction  \nAccurate modeling of radiation transport is essential for nuclear energy and fusion engineering, particle accelerator operation, medical physics, outer space exploration, and advanced manufacturing, which often require analysis and design of systems operating in radiation-rich environments. Monte Carlo (MC) radiation transport codes such as PHITS [1], MCNP [2], GEANT4 [3], and FLUKA [4] are state-of-the-art tools for predicting how neutrons, photons, electrons, and ions interact with materials in arbitrary three-dimensional (3D) geometries. These platforms encode decades of validated physics, but their simulation setup, execution, and validation of results remain challenging even for experienced designers. Constructing a valid simulation requires  \nc Application  \nMedical physics  \n\n|  |\n| --- |\n|  |\n\n150 MeV proton beam, water  \nShielding & facility design  \n662 keV  beam, lead  \nDetection & spectroscopy  \nCs-137 source + NaI(Tl)  \nFigure 1 . PHITS at a glance. (a) A PHITS input file consists of a sequence of named keyword sections defining simulation parameters, particle sources, materials, geometry, and tally requests. (b) During execution, particles are emitted from the specified source, transported through the geometry, and scored according to user-defined [T-*] tally definitions. In the example shown, a neutron source emits neutrons that propagate through the geometry and interact with the material region M1 (orange), producing secondary particles at the interaction points (orange dots); the simulation returns the neutron track-length flux scored by a [T-Track] tally. (c) The same radiation-transport framework underpins diverse application areas; the panels show three representative PHITS simulations from this study—medical physics (a proton beam in water), shielding and facility design (a photon beam in lead), and radiation detection and spectroscopy (a Cs-137 source and a NaI(Tl) detec","cbCaimvROskagUAV","https://ap.wps.com/l/cbCaimvROskagUAV","pdf",2160260,1,21,"English","en",105,"# Introduction\n## PHITS overview and workflow challenges\n## Motivation for AI-assisted radiation-transport input generation\n## Related work and benchmarks","[{\"question\":\"What is PHITSBench designed to evaluate?\",\"answer\":\"PHITSBench evaluates AI systems that generate PHITS radiation-transport inputs from natural language, using execution-scored tasks and observable-level agreement with references.\"},{\"question\":\"How are tasks organized in PHITSBench?\",\"answer\":\"Tasks are grouped into three workflow categories: parameter editing, syntax repair, and complete simulation generation from natural-language descriptions.\"},{\"question\":\"Why do models perform worse on generating simulations from scratch?\",\"answer\":\"Failure analysis indicates remaining errors are dominated more by incorrect selection and configuration of physical observables than by syntax generation, especially on the Reproduce track.\"}]",1784205761,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"phitsbench-an-execution-scored-benchmark-for-ai-assisted-phits-radiation-transport-input-generation-using-natural-language","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/phitsbench-an-execution-scored-benchmark-for-ai-assisted-phits-radiation-transport-input-generation-using-natural-language/85716/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is PHITSBench designed to evaluate?","Question",{"text":75,"@type":76},"PHITSBench evaluates AI systems that generate PHITS radiation-transport inputs from natural language, using execution-scored tasks and observable-level agreement with references.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are tasks organized in PHITSBench?",{"text":80,"@type":76},"Tasks are grouped into three workflow categories: parameter editing, syntax repair, and complete simulation generation from natural-language descriptions.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do models perform worse on generating simulations from scratch?",{"text":84,"@type":76},"Failure analysis indicates remaining errors are dominated more by incorrect selection and configuration of physical observables than by syntax generation, especially on the Reproduce track.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]