[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119731-en":3,"doc-seo-119731-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119731,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",6,"Technology","From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams","A machine learning final exam at leading universities typically requires faculty days to author and students hours to solve. Large language models are shown to pass machine learning finals at a human level on publicly available exam questions after training, and to automatically generate new, high-quality exam questions within seconds. A curated online dataset and benchmark enable answering and question generation, supported by automated checkers for multiple-choice, numeric, and expression answers, with ablation studies comparing prompting and few-shot strategies.","From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams  \narXiv :2206 .05442v7 [ cs .LG] 28 Jun 2023  \nIddo Drori  \nMIT, Columbia University, BU Cambridge, USA [idrori@mit.edu](idrori@mit.edu)  \nSarah Zhang  \nMIT Cambridge, USA [sazhang@mit.edu](sazhang@mit.edu)  \nPedro Lantigua  \nMIT Cambridge, USA[lantigua@mit.edu](lantigua@mit.edu)  \nSarah J. Zhang  \nMIT Cambridge, USA [sjzhang@mit.edu](sjzhang@mit.edu)  \nKeith Tyser  \nBoston University Boston, USA [ktyser@bu.edu](ktyser@bu.edu)  \nSaisamrit Surbehera  \nColumbia University New York, USA [ss6365@columbia.edu](ss6365@columbia.edu)  \nReece Shuttleworth  \nMIT Cambridge, USA[rshuttle@mit.edu](rshuttle@mit.edu)  \nZad Chin  \nHarvard University Cambridge, USA[zadchin@college.harvard.edu](zadchin@college.harvard.edu)  \nGregory Hunter  \nColumbia University New York, USA[geh2129@columbia.edu](geh2129@columbia.edu)  \nDerek Austin Columbia University  \nNew York, USA[da2986@columbia.edu](da2986@columbia.edu)  \nLeonard Tang  \nHarvard University  \nCambridge, USA [leonardtang@college.harvard.edu](leonardtang@college.harvard.edu)  \nYann Hicke Cornell University  \nIthaca, USA[ylh8@cornell.edu](ylh8@cornell.edu)  \nSage Simhon  \nMIT Cambridge, USA [simhon@mit.edu](simhon@mit.edu)  \nSathwik Karnik  \nMIT Cambridge, USA [skarnik@mit.edu](skarnik@mit.edu)  \nDarnell Granberry  \nMIT Cambridge, USA [darnellg@mit.edu](darnellg@mit.edu)  \nMadeleine Udell  \nStanford University Stanford, USA [udell@stanford.edu](udell@stanford.edu)  \nABSTRACT  \nA ﬁnal exam in machine learning at a top institution such as MIT, Harvard, or Cornell typically takes faculty days to write, and students hours to solve. We demonstrate that large language models pass machine learning ﬁnals at a human level, on ﬁnals available online after the models were trained, and automatically generate new human-quality ﬁnal exam questions in seconds. Previous work has developed program synthesis and few-shot learning methods to solve university-level problem set questions in mathematics and STEM courses. In this work, we develop and compare methods that solve ﬁnal exams, which diﬀer from problem sets  \nPermission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for proﬁtor commercial advantage and that copies bear this notice and the full citation on theﬁrst page. Copyrights for third-party components ofthis work must be honored.  \nFor all other uses, contact the owner/author(s) . KDD ’23, August 6–10, 2023, Long Beach, CA, USA © 2023 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-0103-0/23/08 .  \n[https://doi.org/10.1145/3580305.3599827](https://doi.org/10.1145/3580305.3599827)  \nin several ways: the questions are longer, have multiple parts, are more complicated, and span a broader set of topics. We curate adataset and benchmark of questions from machine learning ﬁnal exams available online and code for answering these questions and generating new questions. We show how to generate new questions from other questions and course notes. For reproducibility and future research on this ﬁnal exam benchmark, we use automatic checkers for multiple-choice, numeric, and questions with expression answers. We perform ablation studies comparing zeroshot learning with few-shot learning and chain-of-thought prompting using GPT-3, OPT, Codex, and ChatGPT across machine learning topics and ﬁnd that few-shot learning methods perform best. We highlight the transformative potential of language models to streamline the writing and solution of large-scale assessments, signiﬁcantly reducing the workload from human days to mere machine seconds. Our results suggest that rather than banning large language models such as ChatGPTin class, instructors should teach students to harness them by asking students meta-questions about  \ncorrectness, completeness, and originality of the responses generated","cbCain6HMMhToqgj","https://ap.wps.com/l/cbCain6HMMhToqgj","pdf",186750,1,9,"English","en",105,"# Introduction\n## Dataset and benchmark\n## Automated answering and checkers\n## Few-shot vs zero-shot and prompting ablations\n## Meta-question evaluation","[{\"question\":\"How do the authors make language models answer machine learning final exams?\",\"answer\":\"They curate a structured dataset of online machine learning final exam questions and use automated checkers to evaluate multiple-choice, numeric, and expression-answer formats. The models are prompted to solve and explain, then validated through the benchmark checks.\"},{\"question\":\"What is the key result about model performance on final exams?\",\"answer\":\"The best-performing approaches allow large language models to pass machine learning finals at a human level on available exam questions, while also generating new questions quickly.\"},{\"question\":\"Which learning strategy performs best in the comparisons, and how is it evaluated?\",\"answer\":\"Few-shot learning methods perform best compared with zero-shot learning and other prompting approaches. The evaluation includes meta-questions assessing correctness, completeness, and originality of the generated responses.\"}]","From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams | PDF",1785725998,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-human-days-to-machine-seconds-automatically-answering-and-generating-machine-learning-final-exams","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-human-days-to-machine-seconds-automatically-answering-and-generating-machine-learning-final-exams/119731/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How do the authors make language models answer machine learning final exams?","Question",{"text":75,"@type":76},"They curate a structured dataset of online machine learning final exam questions and use automated checkers to evaluate multiple-choice, numeric, and expression-answer formats. The models are prompted to solve and explain, then validated through the benchmark checks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key result about model performance on final exams?",{"text":80,"@type":76},"The best-performing approaches allow large language models to pass machine learning finals at a human level on available exam questions, while also generating new questions quickly.",{"name":82,"@type":73,"acceptedAnswer":83},"Which learning strategy performs best in the comparisons, and how is it evaluated?",{"text":84,"@type":76},"Few-shot learning methods perform best compared with zero-shot learning and other prompting approaches. The evaluation includes meta-questions assessing correctness, completeness, and originality of the generated responses.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]