[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85948-en":3,"doc-seo-85948-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85948,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Towards Autonomous and Auditable Medical Imaging Model Development","Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by planning, executing code, debugging, and using empirical feedback. Applying such agentic workflows to medical imaging is hindered by modality-specific experimentation needs and stringent validation requirements for metrics and prediction artifacts. AMID introduces an autonomous multi-agent framework that plans data-conditioned method lanes and performs verification-guided two-stage optimization to deliver high-performing, verifiable medical imaging models across 20 challenge tasks.","arXiv :2607 . 10522v1 [ cs .CV] 12 Jul 2026  \nTowards Autonomous and Auditable Medical Imaging Model Development  \nShengyuan Liu 1 ,∗ , Jia-Xuan Jiang 1 ,∗ , Boyun Zheng 1 , Cheng Wang 1 , Zipei Wang2 , Wentao Pan 1 , Hongtao Wu 1 , Houwen Peng5 , Yu Gu3 , Lichao Sun4 ,†, Yixuan Yuan 1 ,†  \n1The Chinese University of Hong Kong, 2Institute of Automation, Chinese Academy of Sciences, 3Microsoft Research, 4Lehigh University, 5Independent Researcher  \n∗Equal contribution, †Corresponding author  \nLarge anguage model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts. Here we introduce AMID, an autonomous multi-agent framework for medical imaging model development. AMID first proposes Data-Conditioned Method Planning, which refines coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources. It then develops Verification-Guided Two-Stage Optimization, moving from broad early exploration of diverse method lanes to selective exploitation of promising candidates while enforcing strict verification of validation protocols, metric computation, and prediction artifacts throughout the optimization. Across 20 medical imaging challenge tasks spanning diverse modalities and prediction types, AMID outperformed evaluated general-purpose MLE systems and, on several tasks, approached or matched strong human-designed challenge solutions. These results suggest that AMID can turn task-specific medical imaging model development from bespoke manual engineering into an agentic workflow for producing high-performing and auditable model artifacts across heterogeneous tasks.  \nNote: This is an ongoing preliminary technical report.  \nProject page: [https://github.com/CUHK-AIM-Group/AMID](https://github.com/CUHK-AIM-Group/AMID)  \nScore  \n0.9  \n0.8  \n0.7  \n0.6  \n0.5  \n\n| Stage 1: Behavior-gated exploration |  |\n| --- | --- |\n| parallel route portfolio\u003Cbr> |  |\n| \u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr> |  |\n|  |  |\n|  |  |\n|  |  |\n\nStage 2: Selective exploitation  \n1-fold → multi-fold   \nfinal submission  \n0 20 40 60 80 100 120  \nAttempt index  \nPerformance on 20 Medical Imaging Challenges  \nAverage metric score  \n1.0  \n0.8  \n0.6  \n0.4  \n0.2  \n0  \n\n|  AIDE  Ours (Codex)\u003Cbr> ML-Master  Human level\u003Cbr> R&D-Agent |  |  |  |  |  |  |  |  |  |  |  |  |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n|  |  |  |  |  |  |  |  |  |  |  |  |  |\n|  |  |  |  |  |  |  |  |  |  |  |  |  |\n|  |  |  |  |  |  |  |  |  |  |  |  |  |\n|  |  |  |  |  |  |  |  |  |  |  |  |  |\n\nDetection Segmentation Average  \nFigure 1 Overview of AMID’s verification-guided optimization and benchmark performance. The two-stage optimization moves from broad exploration of parallel method lanes to selective exploitation of promoted candidates, while reviewer checks verify validation protocols, metric computation, and prediction artifacts throughout the optimization; shaded regions mark the two stages, and the star marks the final selected submission. The category-level bar chart reports average metric scores over 20 ReX-MLE medical imaging challenges.  \n1 Introduction  \nLarge language model (LLM) agents are emerging as a promising paradigm for automating machine learning engineering (MLE), shifting model development from one-shot code generation toward iterative, feedback-driven experimentation (Aygün et al., 2026; Chan et al., 2024; Jimenez et al., 2024) . In this paradigm, agents iteratively transform methodological ideas into executable experiments, using failures, validation signals, and accumulated evidence to guide subsequent refinements. Existing MLE agents (Jiang et al., 2025; Liu et al., 2025; Yang","cbCaia2iqeSndgdx","https://ap.wps.com/l/cbCaia2iqeSndgdx","pdf",2173069,5,1,18,"English","en",105,"# Introduction\n# AMID Framework Overview\n## Stage 1: Behavior-gated exploration\n## Stage 2: Selective exploitation","[{\"question\":\"What is AMID and what problem does it address in medical imaging model development?\",\"answer\":\"AMID is an autonomous multi-agent framework designed to automate medical imaging model development while meeting strict requirements for validation protocols and auditable prediction artifacts.\"},{\"question\":\"How does AMID generate candidate methods for a given medical imaging task?\",\"answer\":\"It first proposes Data-Conditioned Method Planning, refining coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources.\"},{\"question\":\"What role does verification-guided optimization play in AMID?\",\"answer\":\"It enforces strict verification of validation protocols, metric computation, and prediction artifacts throughout optimization, moving from broad exploration of method lanes to selective exploitation of promising candidates.\"}]",1784207316,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"towards-autonomous-and-auditable-medical-imaging-model-development","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-autonomous-and-auditable-medical-imaging-model-development/85948/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is AMID and what problem does it address in medical imaging model development?","Question",{"text":76,"@type":77},"AMID is an autonomous multi-agent framework designed to automate medical imaging model development while meeting strict requirements for validation protocols and auditable prediction artifacts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does AMID generate candidate methods for a given medical imaging task?",{"text":81,"@type":77},"It first proposes Data-Conditioned Method Planning, refining coarse task-level search spaces into executable, parallelizable method lanes grounded in task-specific data analysis and runnable medical-imaging resources.",{"name":83,"@type":74,"acceptedAnswer":84},"What role does verification-guided optimization play in AMID?",{"text":85,"@type":77},"It enforces strict verification of validation protocols, metric computation, and prediction artifacts throughout optimization, moving from broad exploration of method lanes to selective exploitation of promising candidates.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]