[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84988-en":3,"doc-seo-84988-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84988,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Evaluating Static and Process Evidence for Code Authorship in Programming Education","Programming instructors need to judge whether a student submission matches the learner’s prior programming profile, where code similarity alone is often insufficient. Existing source-code authorship methods are typically tested on contest or open-source datasets, but educational repositories differ: students share assignments while their practices are still forming. This study uses task-aware evaluation across six matched educational comparisons to test whether process features visible in repositories add information beyond final code.","arXiv :2607 .07400v 1 [ cs . SE] 8 Jul 2026  \nEvaluating Static and Process Evidence for Code Authorship in Programming Education  \nMAREK HORVÁTH, Department of Computers and Informatics, Technical University of Košice, Slovakia  \nIn programming courses, instructors may need to interpret whether a submission is consistent with a student’s prior programming profile, especially when code similarity alone is inconclusive. Existing source-code authorship methods are often evaluated on programming-contest or open-source datasets, where reusable templates and local code patterns can produce strong author-related signal. Educational repositories present a different setting. Students solve shared assignments while their programming practices are still developing. This study uses task-aware evaluation to contrast these production contexts and tests whether repository-visible process features add information beyond final code in six matched educational comparisons. Contest data provide a high-signal contrast, with a Kick Start mean top-1 of 0.938 . Educational datasets produce substantially lower attribution performance. Adding process features raises the educational mean from 0.094 to 0.233 and mean pairwise verification ROC–AUC from 0.556 to 0.752 . The comparisons show that measured signal depends on production context and that process patterns can complement weak final-code signal in educational repositories. Such models are therefore appropriate only as instructor-mediated decision support, not as independent proof of authorship.  \nCCS Concepts: • Social and professional topics → Computing education; • Software and its engineering → Empirical software validation; • Applied computing → Interactive learning environments; • Computing methodologies → Supervised learning by classification.  \nAdditional Key Words and Phrases: source code authorship attribution, authorship verification, programming education, academic integrity, repository-visible process features, task-aware evaluation, machine learning  \nACM Reference Format:  \nMarek Horváth. 2026. Evaluating Static and Process Evidence for Code Authorship in Programming Education. ACM Trans. Comput. Educ. 1, 1 (July 2026), 31 pages. [https://doi.org/10.1145/nnnnnnn.nnnnnnn](https://doi.org/10.1145/nnnnnnn.nnnnnnn)  \n1 Introduction  \nSource code is a formal artifact. It must be accepted by a compiler or interpreter, express an executable procedure, and often satisfy automated tests. These constraints might appear to leave little room for individual expression. Yet the same task can usually be implemented in several ways. Programmers choose how to decompose a solution, name variables and functions, use comments, organize control flow, reuse idioms, and revise a solution over time. Repeated choices of this kind form a programming style that can be observed, measured, and modeled.  \nSource code authorship analysis assumes that some part of this style remains stable across an author’s work. That assumption is useful but conditional. Programming style is shaped by language, problem type, libraries, development environment, examples, templates, team conventions, and experience. Programmers can learn, imitate, adapt, or intentionally change their style. This instability is especially visible in education, where students are still acquiring programming habits and work under shared instructional constraints. Authorship analysis therefore cannot be reduced to selecting the classifier with the highest accuracy.  \nThe literature on source code authorship has produced strong results in controlled settings, especially on programming-contest datasets such as Google Code Jam and on other large code collections [Abuhamad et al. 2021, 2019; Caliskan-Islam et al. 2015; White and Sprague 2021] . These datasets are valuable because they contain many authors, many comparable tasks, and enough samples for training machine-learning models. They also have a particular character. Contest  \nAuthor’s Contact Infor","cbCaink32VV3DuZm","https://ap.wps.com/l/cbCaink32VV3DuZm","pdf",612460,3,1,31,"English","en",105,"# Introduction\n## Source code as an observable artifact\n## Limits of closed-world authorship assumptions\n## Educational repositories and task leakage","[{\"question\":\"Why is code-similarity alone often inconclusive for code authorship in programming education?\",\"answer\":\"Educational repositories impose shared instructional and infrastructure constraints, so the strongest patterns in the data may reflect the task or assignment structure rather than individual author style.\"},{\"question\":\"What does the study compare between different data production contexts?\",\"answer\":\"It contrasts contest-like production contexts with educational repository contexts, using task-aware evaluation to test whether repository-visible process features provide additional information beyond final code.\"},{\"question\":\"How do process features affect authorship attribution performance in educational datasets?\",\"answer\":\"Adding process features increases educational mean attribution performance from 0.094 to 0.233 and improves mean pairwise verification ROC–AUC from 0.556 to 0.752.\"}]",1784200064,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-static-and-process-evidence-for-code-authorship-in-programming-education","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evaluating-static-and-process-evidence-for-code-authorship-in-programming-education/84988/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is code-similarity alone often inconclusive for code authorship in programming education?","Question",{"text":75,"@type":76},"Educational repositories impose shared instructional and infrastructure constraints, so the strongest patterns in the data may reflect the task or assignment structure rather than individual author style.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the study compare between different data production contexts?",{"text":80,"@type":76},"It contrasts contest-like production contexts with educational repository contexts, using task-aware evaluation to test whether repository-visible process features provide additional information beyond final code.",{"name":82,"@type":73,"acceptedAnswer":83},"How do process features affect authorship attribution performance in educational datasets?",{"text":84,"@type":76},"Adding process features increases educational mean attribution performance from 0.094 to 0.233 and improves mean pairwise verification ROC–AUC from 0.556 to 0.752.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]