[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119411-en":3,"doc-seo-119411-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119411,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","HARDML - A Benchmark for Evaluating Data Science and Machine Learning Knowledge and Reasoning in AI","HardML is a benchmark created to measure knowledge and reasoning in data science and machine learning through 100 challenging multiple-choice questions. The items are handcrafted over six months with strong originality to reduce benchmark data contamination. Even senior machine learning engineers may find many questions difficult. Current state-of-the-art AI models reach about a 30% error rate, roughly triple that on MMLU-ML, making HardML a rigorous, modern testbed to track progress of advanced AI models.","STUDIA UNIV. BABES¸–BOLYAI, INFORMATICA, Volume LXIX, Number 2, 2024 DOI: 10.24193/subbi.2024.2.04  \nHARDML: A BENCHMARK FOR EVALUATING DATA SCIENCE AND MACHINE LEARNING KNOWLEDGE AND  \nREASONING IN AI  \nTIDOR-VLAD PRICOPE  \nAbstract. We present HardML, a benchmark designed to evaluate the knowledge and reasoning abilities in the fields of data science and machine learning. HardML comprises a diverse set of 100 challenging multiplechoice questions, handcrafted over a period of 6 months, covering the most popular and modern branches of data science and machine learning. These questions are challenging even for a typical Senior Machine Learning Engineer to answer correctly. To minimize the risk of data contamination, HardML uses mostly original content devised by the author. Current stateof-the-art AI models achieve a 30% error rate on this benchmark, which isabout 3 times larger than the one achieved on the equivalent, well-known MMLU-ML. While HardML is limited in scope and not aiming to push the frontier—primarily due to its multiple-choice nature—it serves as a rigorous and modern testbed to quantify and track the progress of top AI.  \nWhile plenty benchmarks and experimentation in LLM evaluation exist in other STEM fields like mathematics, physics and chemistry, the sub-fields of data science and machine learning remain fairly underexplored.  \n1. Introduction  \nRecent advancements in large language models (LLMs) have led to significant progress in natural language processing tasks such as translation, summarization, question answering, and code generation [1, 2] . These models have been extensively evaluated using benchmarks covering a wide range of subjects, providing valuable insights into their capabilities [3, 4] . For instance, the Massive Multitask Language Understanding (MMLU) benchmark  \nReceived by the editors: 22 January 2025 .  \n2020 Mathematics Subject Classification. 68T50, 68T07, 68T05, 68T20 .  \nKey words and phrases. Large Language Models, Machine Learning Education, Multiple Choice Benchmark, NLP Benchmarks, Evaluation of AI Systems.  \n© Studia UBB Informatica. Published by Babe¸s-Bolyai University  \n This work is licensed under a Creative Commons Attribution-NonCommercialNoDerivatives 4 .0 International Licence.  \n60 TIDOR-VLAD PRICOPE  \n[5] assesses LLMs across diverse disciplines, including STEM fields like mathematics, physics, and chemistry [6, 7, 8] . However, data science (DS) and machine learning (ML) have received relatively little attention in benchmarking efforts. The MMLU test set contains only 112 machine learning questions. Moreover, in the few instances where these domains have been explored, stateof-the-art AI models achieve near-saturation performance, rendering existing benchmarks less effective for distinguishing model capabilities.  \nIt is imperative to devise novel benchmarks that keep up with the rapid advancements in LLMs. This necessity is exemplified in the FrontierMath benchmark [15], which introduces a future-proof evaluation for mathematics by presenting problems that remain unsolved by over 98% of current AI models. Such benchmarks are crucial for continuing to challenge and develop advanced AI systems.  \nData science and machine learning are foundational to modern artificial intelligence, driving advancements in everything from healthcare to finance [9, 10] . Mastery in these fields requires not only theoretical understanding but also practical skills in applying algorithms, statistical methods, and computational techniques to solve complex, real-world problems [11, 12] . As AI systems become increasingly involved in DS and ML tasks—ranging from automated model training to data analysis—it is crucial to assess their proficiency and reasoning abilities in these areas. However, as of January 2025, benchmarks in this domain are very limited. The most notable examples include the test ML subsection of MMLU (MMLU-ML) [5], which consists of 112 multiple-choice questions, and MLE-benc","cbCaiulldFuEh2YD","https://ap.wps.com/l/cbCaiulldFuEh2YD","pdf",987525,1,18,"English","en",105,"# Introduction\n## Motivation for new DS/ML benchmarks\n## HardML design and coverage\n## Originality and contamination control","[{\"question\":\"What does HardML evaluate in AI models?\",\"answer\":\"HardML evaluates knowledge and reasoning abilities specifically for data science and machine learning domains.\"},{\"question\":\"How is HardML constructed to reduce data contamination risk?\",\"answer\":\"It uses mostly original question content created by the author, minimizing the chance that models were trained on benchmark items.\"},{\"question\":\"What is the format and scale of the HardML benchmark?\",\"answer\":\"HardML consists of 100 multiple-choice questions using an MMLU-like framework, with an added characteristic that more than one answer can be correct.\"}]","HARDML - A Benchmark for Evaluating Data Science and Machine Learning Knowledge and Reasoning in AI | PDF",1785724156,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hardml-a-benchmark-for-evaluating-data-science-and-machine-learning-knowledge-and-reasoning-in-ai","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hardml-a-benchmark-for-evaluating-data-science-and-machine-learning-knowledge-and-reasoning-in-ai/119411/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does HardML evaluate in AI models?","Question",{"text":75,"@type":76},"HardML evaluates knowledge and reasoning abilities specifically for data science and machine learning domains.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is HardML constructed to reduce data contamination risk?",{"text":80,"@type":76},"It uses mostly original question content created by the author, minimizing the chance that models were trained on benchmark items.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the format and scale of the HardML benchmark?",{"text":84,"@type":76},"HardML consists of 100 multiple-choice questions using an MMLU-like framework, with an added characteristic that more than one answer can be correct.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]