[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86037-en":3,"doc-seo-86037-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86037,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Imaging-101 Benchmarking LLM Coding Agents on Scientific Computational Imaging","Imaging-101 provides a benchmark for computational imaging by standardizing 57 expert-verified tasks across six scientific domains into a four-stage pipeline: preprocessing, forward physics modeling, inverse solver, and visualization. Three evaluation tracks—planning, function-level unit tests, and end-to-end reconstruction—assess distinct agent capabilities from code preparation to verified reconstruction. Testing seven frontier LLMs reveals systematic difficulties in agentic coding for computational imaging, including algorithm selection, physical convention handling, and pipeline integration, motivating domain-specialized, skill-augmented agents.","Imaging-101: Benchmarking LLM Coding Agentson Scientific Computational Imaging  \nSiyi Chen∗ , Jiahe Ying∗ , Yixuan Jia∗ , Yuxuan Gu∗ , Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren,  \nBinhong Gao, Yubing Li, Member, IEEE, Tianhan Zhang, and He Sun†, Member, IEEE  \nAbstract—Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific disciplines, yet building a correct reconstruction pipeline demands deep domain expertise and remains laborious even for domain scientists. We introduce Imaging-101, a benchmark of 57 expert-verified computational imaging tasks spanning six scientific domains, each grounded in a peer-reviewed paper and canonicalized into a standardized four-stage pipeline (preprocessing, forward physics modeling, inverse solver, and visualization) . Three evaluation tracks (planning, function-level unit tests, and end-to-end reconstruction) probe distinct agent capabilities across the full pipeline. Evaluating seven frontier LLMs uncovers systematic challenges in applying coding agents to computational imaging that go beyond those exposed by general coding benchmarks, spanning algorithm selection, physical convention handling, and pipeline integration. These findings highlight concrete capability gaps and point towardskill-augmented, domain-specialized agents as a practical path to reliable computational imaging assistance.  \nIndex Terms—Scientific Computational Imaging, Inverse Problems, Large Language Models, Scientific Benchmarks, Agentic Systems  \n~~ ~~ ✦ ~~ ~~  \narXiv :2607 . 10789v 1 [ cs .AI] 12 Jul 2026  \n1 INTRODUCTION  \nScience begins where measurement begins. Modern science relies fundamentally on computational imaging, the recovery of hidden signals from indirect, noisy, and incomplete measurements, to infer the unknown states of physical, chemical, and biological systems [1], [2] . Mathematically, each such task is an inverse problem: while a forward model predicts observations from known parameters, the imaging pipeline operates in the opposite direction, recovering latent quantities that cannot be measured directly. This framework underpins quantitative discovery across scientific disciplines [3], [4], [5], [6], from reconstructing black hole images [7], mapping subsurface geology [8] to recovering anatomical structure in CT [9],[10],[11] and MRI [12],[13],[14] .  \nDespite their ubiquity, solving a single computational imaging task in practice is laborious [15], [16] . Real measurements are noisy, sparse, and often poorly conditioned,  \n• This work was supported by the Shanghai Municipal Science and Technology Major Project (2025SHZDZX026D03), the Natural Science Foundation of Beijing Municipality (Z240010) and the National Natural Science Foundation of China (32450631 and 62371007).  \n• † [Corresponding author: He Sun. E-mail: hesun@pku.edu.cn](Corresponding author: He Sun. E-mail: hesun@pku.edu.cn).  \n• ∗ Siyi Chen, Jiahe Ying, Yixuan Jia, and Yuxuan Gu contributed equally to this work.  \n• Siyi Chen, Jiahe Ying, Yuxuan Gu, Enze Ye, Weimin Bai, Zhijun Zeng, Shaochi Ren, Binhong Gao, and He Sun are with the College of Future Technology and the National Biomedical Imaging Center, Peking University, Beijing, 100871, China.  \n• Jiahe Ying is also with the AI for Science Institute (AISI), Beijing, 100080, China.  \n• Yixuan Jia is with the University of Michigan, Ann Arbor, MI 48105, United States.  \n• Yubing Li is with the State Key Laboratory of Acoustics and Marine Information, Institute of Acoustics, Chinese Academy of Sciences, Beijing, 100190, China, and also with the University of Chinese Academy of Sciences, Beijing, 100049, China.  \n• Tianhan Zhang is with the School of Astronautics, Beihang University, Beijing, 100191, China. Tianhan Zhang is also with the Key Laboratory of Spacecraft Design Optimization and Dynamic Simulation Technologies, Ministry of Education, Beijing, 102206, China.  \nleading to severe ill-posedne","cbCaipH7SQdIf7Ev","https://ap.wps.com/l/cbCaipH7SQdIf7Ev","pdf",6428208,4,1,45,"English","en",105,"# Introduction\n## Computational imaging and inverse problems\n## Motivation for LLM coding agents\n# Imaging-101 Benchmark\n## Task canonicalization into a four-stage pipeline\n## Evaluation tracks and grading protocol\n# Findings","[{\"question\":\"What is Imaging-101 and what does it benchmark?\",\"answer\":\"Imaging-101 benchmarks coding agents on 57 expert-verified computational imaging tasks spanning six scientific domains. Each task is grounded in peer-reviewed literature and evaluated through executable-code grading with numerical tolerance checks.\"},{\"question\":\"How are the computational imaging tasks standardized in the benchmark?\",\"answer\":\"Every task is canonicalized into a standardized four-stage pipeline: preprocessing, forward physics modeling, inverse solver, and visualization. This lets agents use a consistent interface across domains.\"},{\"question\":\"How are agent capabilities evaluated in Imaging-101?\",\"answer\":\"Evaluation uses three tracks: planning, function-level unit tests, and end-to-end reconstruction. These tracks test different competencies across the full pipeline rather than only partial workflow steps.\"}]",1784207990,113,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"imaging-101-benchmarking-llm-coding-agents-on-scientific-computational-imaging","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/imaging-101-benchmarking-llm-coding-agents-on-scientific-computational-imaging/86037/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Imaging-101 and what does it benchmark?","Question",{"text":75,"@type":76},"Imaging-101 benchmarks coding agents on 57 expert-verified computational imaging tasks spanning six scientific domains. Each task is grounded in peer-reviewed literature and evaluated through executable-code grading with numerical tolerance checks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the computational imaging tasks standardized in the benchmark?",{"text":80,"@type":76},"Every task is canonicalized into a standardized four-stage pipeline: preprocessing, forward physics modeling, inverse solver, and visualization. This lets agents use a consistent interface across domains.",{"name":82,"@type":73,"acceptedAnswer":83},"How are agent capabilities evaluated in Imaging-101?",{"text":84,"@type":76},"Evaluation uses three tracks: planning, function-level unit tests, and end-to-end reconstruction. These tracks test different competencies across the full pipeline rather than only partial workflow steps.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]