[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122036-en":3,"doc-seo-122036-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122036,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","David vs. Goliath - comparing conventional machine learning and a large language model for assessing students’ concept use in a physics problem","Large language models offer new opportunities for educational research, especially for assessment, yet they face limitations such as hallucinating knowledge, weak explainability of predictions, and high resource expenditure. Conventional machine learning may be preferable when researchers need tighter control over modeling decisions. This study evaluates how conventional machine learning algorithms compare with a recently advanced large language model for assessing students’ concept use in a physics problem-solving task. Results show conventional machine learning outperforms the large language model, with analyses of classification behavior supporting the conclusion that both approaches can be complementary when labeled data is available.","TYPE Original Research PUBLISHED 18 September 2024 DOI 10. 3389/frai.2024.1408817  \nOPEN ACCESS  \nEDITED BY  \nGeorgios Leontidis,  \nUniversity of Aberdeen, United Kingdom  \nREVIEWED BY  \nMichael Flor,  \nEducational Testing Service, United States Antonio Sarasa-Cabezuelo, Complutense University of Madrid, Spain  \n*CORRESPONDENCE  \nFabian Kieser  \n kieser@ph-heidelberg.de  \nRECEIVED 15 April 2024  \nACCEPTED 30 August 2024  \nPUBLISHED 18 September 2024  \nCITATION  \nKieser F, Tschisgale P, Rauh S, Bai X, Maus H, Petersen S, Stede M, Neumann K and Wul􀀀 P (2024) David vs. Goliath: comparing conventional machine learning and a large language model for assessing students’concept use in a physics problem.  \nFront. Artif. Intell. 7:1408817 .  \ndoi: 10.3389/frai.2024.1408817  \nCOPYRIGHT  \n© 2024 Kieser, Tschisgale, Rauh, Bai, Maus, Petersen, Stede, Neumann and Wul􀀀 . This isan open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nDavid vs. Goliath: comparing conventional machine learning and a large language model for assessing students’ concept use in a physics problem  \nFabian Kieser1*, Paul Tschisgale2 , Sophia Rauh3 , Xiaoyu Bai3 , Holger Maus2 , Stefan Petersen2 , Manfred Stede3 ,  \nKnut Neumann2 and Peter Wul􀀀1  \n1 Physics and Physics Education Research, Heidelberg University of Education, Heidelberg, Germany,  \n2 Department of Physics Education, Leibniz Institute for Science and Mathematics Education, Kiel, Germany, 3Applied Computational Linguistics, University of Potsdam, Potsdam, Germany  \nLarge language models have been shown to excel in many di􀀀erent tasks across disciplines and research sites. They provide novel opportunities to enhance educational research and instruction in di􀀀erent ways such as assessment. However, these methods have also been shown to have fundamental limitations. These relate, among others, to hallucinating knowledge, explainability of model decisions, and resource expenditure. As such, more conventional machine learning algorithms might be more convenient for speciﬁc research problems because they allow researchers more control over their research. Yet, the circumstances in which either conventional machine learning or large language models are preferable choices are not well understood. This study seeks to answer the question to what extent either conventional machine learning algorithms or a recently advanced large language model performs better in assessing students’ concept use in a physics problem-solving task. We found that conventional machine learning algorithms in combination outperformed the large language model. Model decisions were then analyzed via closer examination of the models’ classiﬁcations. We conclude that in speciﬁc contexts, conventional machine learning can supplement large language models, especially when labeled data is available.  \nKEYWORDS  \nlarge language models, machine learning, natural language processing, problem solving, explainable AI  \n1 Introduction  \nThe introduction of ChatGPT, a conversational arti􀀂cial intelligence (AI)-based bot, to the public in November 2022 directed attention to large language models (LLMs) . As of 2023, ChatGPT is based on a LLM called Generative Pre-trained Transformer (GPT; versions 3.5, 4, 4V, or 4o) and has proven to perform surprisingly well on a wide range of di􀀓erent tasks in various disciplines—including medicine, law, economics, mathematics, chemistry, and physics (Hallal et al., 2023; West, 2023; Surameery and Shakor, 2023; Sinha et al., 2023) . A number of tasks in education (research) can be tackled using LLMs in general, or ChatGPT more speci􀀂cally","cbCaioZjEKFICLdM","https://ap.wps.com/l/cbCaioZjEKFICLdM","pdf",1788097,1,18,"English","en",105,"# Introduction\n## Opportunities of large language models in education\n## Challenges: hallucination, explainability, bias, and resource use\n# Study focus and comparison goal","[{\"question\":\"What question does the study aim to answer?\",\"answer\":\"It examines to what extent conventional machine learning algorithms or a recently advanced large language model performs better for assessing students’ concept use in a physics problem-solving task.\"},{\"question\":\"What were the main findings from the comparison?\",\"answer\":\"Conventional machine learning algorithms, when combined, outperformed the large language model for the targeted assessment task.\"},{\"question\":\"Why might conventional machine learning be preferred in some educational research contexts?\",\"answer\":\"Conventional methods can provide more control over research design, and they may supplement large language models particularly when labeled data is available, addressing concerns such as hallucination and lack of explainability.\"}]","David vs. Goliath - comparing conventional machine learning and a large language model for assessing students’ concept use in a physics problem | PDF",1785808456,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"david-vs-goliath-comparing-conventional-machine-learning-and-a-large-language-model-for-assessing-students-concept-use-in-a-physics-problem","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/david-vs-goliath-comparing-conventional-machine-learning-and-a-large-language-model-for-assessing-students-concept-use-in-a-physics-problem/122036/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the study aim to answer?","Question",{"text":75,"@type":76},"It examines to what extent conventional machine learning algorithms or a recently advanced large language model performs better for assessing students’ concept use in a physics problem-solving task.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What were the main findings from the comparison?",{"text":80,"@type":76},"Conventional machine learning algorithms, when combined, outperformed the large language model for the targeted assessment task.",{"name":82,"@type":73,"acceptedAnswer":83},"Why might conventional machine learning be preferred in some educational research contexts?",{"text":84,"@type":76},"Conventional methods can provide more control over research design, and they may supplement large language models particularly when labeled data is available, addressing concerns such as hallucination and lack of explainability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]