[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120094-en":3,"doc-seo-120094-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120094,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","MLAgentBench - Evaluating Language Agents on Machine Learning Experimentation - Research paper","MLAgentBench evaluates whether language-model-driven agents can perform end-to-end machine learning experimentation: specifying tasks, modifying code and files, executing Python, and using observed outputs to improve results. The benchmark spans 13 tasks, from boosting CIFAR-10 model accuracy to newer research and Kaggle-style challenges. Agents are instantiated with a ReAct-style framework and benchmarked across multiple LLMs. Results show Claude v3 Opus achieves the highest success rate (37.5% average), with plans and actions that are highly interpretable. Performance varies widely, motivating focus on challenges like long-term planning and hallucination reduction.","MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation  \nQian Huang 1 Jian Vora 1 Percy Liang 1 Jure Leskovec 1  \narXiv :2310 .03302v2 [ cs .LG] 14 Apr 2024  \nAbstract  \nA central aspect of machine learning research is experimentation, the process of designing and running experiments, analyzing the results, and  \niterating towards some positive outcome (e.g., improving accuracy) . Could agents driven by powerful language models perform machine learning experimentation effectively? To answer this question, we introduce MLAgentBench, a suite of 13 tasks ranging from improving model performance on CIFAR-10 to recent research problems like BabyLM. For each task, an agent can perform actions like reading/writing files, executing code, and inspecting outputs. We then construct an agent that can perform ML experimentation based on ReAct framework. We benchmark agents based on Claude v1.0, Claude v2.1, Claude v3 Opus, GPT-4, GPT-4-turbo, Gemini-Pro, and Mixtral and find that a Claude v3 Opus agent is the best in terms of success rate. It can build compelling ML models over many tasks in MLAgentBench with 37.5% average success rate. Our agents also display highly interpretable plans and actions. However, the success rates vary considerably; they span from 100% on well-established older datasets to as low as 0% on recent Kaggle challenges created potentially after the underlying LM was trained. Finally, we identify several key challenges for LM-based agents such as long-term planning and reducing hallucination. 2  \n1. Introduction  \nMuch of the progress in machine learning is driven by effective experimentation: Given a task (e.g., image classification), a researcher develops a method (e.g., choice of model architecture and learning algorithm), runs an experiment, and then evaluates the results. Based on the outcome of  \n1 Stanford University. Correspondence to: Qian Huang \u003Cqh[wang@cs.stanford.edu](wang@cs.stanford.edu) >.  \n2Our code is released at [https://github](https://github.com/snap)[.](https://github.com/snap)[com/snap](https://github.com/snap)stanford/MLAgentBench/ .  \nthe experiment (e.g., validation accuracy), they revise their method to improve performance on the task. This iterative process is challenging, as it requires the researcher to possess extensive prior knowledge about potential methods, to produce functional code, and to interpret experimental results for future improvements.  \nThe complexity and expertise required for successful machine learning experimentation pose significant barriers to entry. In light of these challenges, there has been interest in the possibility of automating aspects of the machine learning workflow, such as Neural Architecture Search (Elskenet al., 2019) and AutoML (He et al., 2021) . The emergence of advanced language models, with their ability to understand and generate human-like text, presents an promising opportunity to further automate ML experimentation end to end. Can we develop an agent capable of conducting machine learning experimentation autonomously?  \nIn this paper, we propose MLAgentBench, the first benchmark for evaluating agents capable of machine learning experimentation (Figure 1) . MLAgentBench is a general framework for specifying experimentation tasks with clear goals and automatically evaluates agents on these tasks. Concretely, each task is specified with a task description, a set of starter files (including starter code and data, e.g., Kaggle data package), and an evaluator that can assign a performance metric score to a final submission (such as test set accuracy of the submitted test set prediction) . Given these, an agent can perform actions like reading/writing files and executing Python code in a workspace. During the agent’s interaction with the environment, we collect its interaction trace for evaluation, which is the agent actions and intermediate snapshots of the workspace (i.e., the set of files and directories in the working directo","cbCaimsVhCF0LRQI","https://ap.wps.com/l/cbCaimsVhCF0LRQI","pdf",1095863,1,39,"English","en",105,"# Abstract\n# Introduction\n## Problem motivation and automation interest\n## Proposed benchmark: MLAgentBench\n## Task specification and evaluation criteria","[{\"question\":\"What is MLAgentBench designed to evaluate?\",\"answer\":\"MLAgentBench evaluates agents’ ability to conduct machine learning experimentation autonomously, including improving task metrics through code and file interactions and producing measurable final outputs.\"},{\"question\":\"How are agent tasks specified and evaluated in MLAgentBench?\",\"answer\":\"Each task includes a task description, starter files (code and data), and an evaluator that scores a final submission using a performance metric such as test accuracy. Agents are also assessed on competence improvement (e.g., +10% over baseline) and efficiency (time and token usage).\"},{\"question\":\"Which language model performed best and what were the key findings?\",\"answer\":\"A Claude v3 Opus agent achieved the best success rate, with 37.5% average success across MLAgentBench tasks. The study also finds that success rates vary significantly by dataset recency, and it highlights challenges such as long-term planning and hallucination reduction.\"}]","MLAgentBench - Evaluating Language Agents on Machine Learning Experimentation - Research paper | PDF",1785728129,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mlagentbench-evaluating-language-agents-on-machine-learning-experimentation-research-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mlagentbench-evaluating-language-agents-on-machine-learning-experimentation-research-paper/120094/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is MLAgentBench designed to evaluate?","Question",{"text":75,"@type":76},"MLAgentBench evaluates agents’ ability to conduct machine learning experimentation autonomously, including improving task metrics through code and file interactions and producing measurable final outputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are agent tasks specified and evaluated in MLAgentBench?",{"text":80,"@type":76},"Each task includes a task description, starter files (code and data), and an evaluator that scores a final submission using a performance metric such as test accuracy. Agents are also assessed on competence improvement (e.g., +10% over baseline) and efficiency (time and token usage).",{"name":82,"@type":73,"acceptedAnswer":83},"Which language model performed best and what were the key findings?",{"text":84,"@type":76},"A Claude v3 Opus agent achieved the best success rate, with 37.5% average success across MLAgentBench tasks. The study also finds that success rates vary significantly by dataset recency, and it highlights challenges such as long-term planning and hallucination reduction.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]