[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117761-en":3,"doc-seo-117761-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117761,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",6,"Technology","Use of Machine Learning Methods in Automatic Assessment Programming Assignments","Programming is widely taught in both traditional and online formats, which increases instructors’ workload for evaluating student submissions. Unit testing supports automation but cannot reliably judge code structure, style, partially correct solutions, or levels of achievement. This thesis investigates machine learning techniques to assess code correctness and quality, aiming to assist grading. Using nine machine learning algorithms on feature sets derived from 500,000+ student submissions, results show automated prediction of assessment-relevant qualities. Findings support using source-code token features for program correctness evaluation and enabling multi-valued quality grading beyond binary pass/fail unit tests.","Technological University Dublin  \nARROW@TU Dublin  \n\n| Masters | Science |\n| --- | --- |\n| 2023\u003Cbr>Use of Machine Learning Methods in Automatic Assessment Programming Assignments\u003Cbr>Botond Tarcsay\u003Cbr>Technological University Dublin, [t.botond@yahoo.co.uk](t.botond@yahoo.co.uk)\u003Cbr>Follow this and additional works at: [https://arrow.tudublin.ie/scienmas](https://arrow.tudublin.ie/scienmas)\u003Cbr> Part of the Computer Sciences Commons |  |\n\nRecommended Citation  \nTarcsay, B. (2023) . Use of Machine Learning Methods in Automatic Assessment Programming Assignments. Technological University Dublin. DOI: 10.21427/EQW1-3S76  \nThis Theses, Masters is brought to you for free and open access by the Science at ARROW@TU Dublin. It has been accepted for inclusion in Masters by an authorized administrator of ARROW@TU Dublin. For more information, please contact [arrow.admin@tudublin.ie](arrow.admin@tudublin.ie), [aisling.coyne@tudublin.ie](aisling.coyne@tudublin.ie), [gerard.connolly@tudublin.ie](gerard.connolly@tudublin.ie).  \nThis work is licensed under a Creative Commons Attribution-Noncommercial-Share Alike 4.0 License  \nUse of Machine Learning Methods in Automatic Assessment Programming  \nAssignments  \nBotond Tarcsay  \nSupervisor: Jelena Vasi Fernando Perez Tellez  \nTU Dublin  \nThis dissertation is submitted for the degree of Master of Philosophy  \nJanuary 2023  \nDeclaration  \nI certify that this thesis which I now submit for examination for the award of Masters by Research, is entirely my own work and has not been taken from the work of others save and to the extent that such work has been cited and acknowledged within the test of my work.  \nThis thesis was prepared according to the regulations for graduate study by research of the Technological University Dublin and has not been submitted in whole or in part for another award in any other third level Institution or University.  \nThe work reported on in this thesis conforms to the principles and requirements of the TU Dublin’s guidelines for ethics in research.  \nTU Dublin has permission to keep, lend or copy this thesis in whole or in part, on condition that any such use of the material of the thesis be duly acknowledged.  \nSigned:  \nBotond Tarcsay January 2023  \nAcknowledgements  \nI would like to express my sincere thanks to Jelena Vasi and Fernando Perez Tellez for their continuous help and hard work throughout the past two years. Their unparalleled guidance and assistance were crucial to this Thesis.  \nI would also like to thank all the TUD lecturers for their work, along with Dr. Stephen Blott for the platform explanation, David Azcona and Prof. Alan Smeaton for the dataset and Keith Quille for the brainstorming sessions.  \nFinally, I would like to thank my partner for her patience and support.  \nAbstract  \nProgramming has become an important skill in today’s world and is taught widely both in traditional settings and online. Instructors need to assess increasing amounts of student work. Unit testing can contribute to the automation of the grading process; however, it cannot assess the structures, style and partially correct source code or differentiate between levels of achievement. The topic of this thesis is an investigation into the use of machine learning methods for assessing the correctness and quality of code, with the ultimate goal of assisting instructors in the grading process. In this research, we have used nine different machine learning algorithms, applied to three distinct types of feature sets, created from over five hundred thousand student code submissions. Prediction scores for some of the models show that the content of the submissions can be assessed in an automated manner. Along with unit testing, this approach has the potential to give instructors a source code-based automated way of assigning more finely differentiated grades than is possible by unit testing alone. This dissertation reports on several findings that confirm the validity of using machine learnin","cbCaimRad2mY3Fy5","https://ap.wps.com/l/cbCaimRad2mY3Fy5","pdf",2877867,1,126,"English","en",105,"# 1 Introduction\n## 1.1 Research Question\n## 1.2 Research Objectives\n## 1.3 Scope and Limitations\n## 1.4 Organization of the Thesis\n# 2 Literature Review\n## 2.1 Manual Grading\n## 2.2 Automated Grading and Feedback Tools\n## 2.3 Code Evaluation for Grading\n### 2.3.1 Semantic Comparison\n### 2.3.2 Unit Testing\n## 2.4 Code Evaluation for Other Purposes\n## 2.5 Summary\n# 3 Methodology\n## 3.1 Data\n### 3.1.1 Raw Data and ByteCode\n### 3.1.2 Derived Datasets\n### 3.1.3 Preliminary Data Analysis\n## 3.2 Models\n### 3.2.1 Models for Line Count Data\n### 3.2.2 Models for Token Count Data\n### 3.2.3 Models for Token Sequence Data\n### 3.2.4 Models for Multi-Class Label Data","[{\"question\":\"Why does unit testing alone not meet programming assignment grading needs?\",\"answer\":\"Unit testing can automate grading to an extent, but it cannot assess code structure and style, handle partially correct source code well, or distinguish achievement levels the same way across grading schemes.\"},{\"question\":\"What is the main goal of this thesis?\",\"answer\":\"To investigate machine learning methods that evaluate the correctness and quality of code so instructors can grade more effectively using automated, code-based predictions.\"},{\"question\":\"How was the machine learning approach evaluated in the research?\",\"answer\":\"Nine machine learning algorithms were applied to three types of feature sets created from 500,000+ student code submissions, and prediction scores were analyzed to verify automated assessment of submission content.\"}]","Use of Machine Learning Methods in Automatic Assessment Programming Assignments | PDF",1785679431,318,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"use-of-machine-learning-methods-in-automatic-assessment-programming-assignments","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/use-of-machine-learning-methods-in-automatic-assessment-programming-assignments/117761/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does unit testing alone not meet programming assignment grading needs?","Question",{"text":75,"@type":76},"Unit testing can automate grading to an extent, but it cannot assess code structure and style, handle partially correct source code well, or distinguish achievement levels the same way across grading schemes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main goal of this thesis?",{"text":80,"@type":76},"To investigate machine learning methods that evaluate the correctness and quality of code so instructors can grade more effectively using automated, code-based predictions.",{"name":82,"@type":73,"acceptedAnswer":83},"How was the machine learning approach evaluated in the research?",{"text":84,"@type":76},"Nine machine learning algorithms were applied to three types of feature sets created from 500,000+ student code submissions, and prediction scores were analyzed to verify automated assessment of submission content.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,113,118,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":111,"slug":112},50,"technology",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},8,"Research & Report",30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]