[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85155-en":3,"doc-seo-85155-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85155,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","When LLM Tutoring Responses Work: Evidence from Student Programming Conversations","As computer science students increasingly rely on LLM tutors, the key question is which response style helps learners continue productively during programming help-seeking. This paper analyzes StudyChat, a public dataset of student–ChatGPT tutoring conversations from an AI course, transformed into 16,851 assistant-response interactions. Local annotation with Gemma 4 labels help-seeking situations, student state, assistant style, and next-turn outcomes. Human validation reaches 82% agreement. Response style significantly predicts productive and unresolved continuation, with verification feedback highest and direct answers lowest. Patterns vary by context.","When LLM Tutoring Responses Work: Evidence from Student  \nProgramming Conversations  \nMohammad Fahim Abrar  \n[fahim@udel.edu](fahim@udel.edu)[ ](fahim@udel.edu)University of Delaware Newark, Delaware, USA  \nShyala Sharmin  \n[shayla@udel.edu](shayla@udel.edu)[ ](shayla@udel.edu)University of Delaware Newark, Delaware, USA  \nRoghayeh Leila Barmaki  \n[rlb@udel.edu](rlb@udel.edu)[ ](rlb@udel.edu)University of Delaware Newark, Delaware, USA  \narXiv :2607 .09919v1 [ cs .HC] 10 Jul 2026  \nFigure 1: Overview of the analysis pipeline. Authentic student-ChatGPT tutoring conversations from StudyChat were transformed into assistant-response interactions, annotated with Gemma 4 for student help-seeking situation, student state, assistant response style, and next-turn outcome, and analyzed to identify response patterns in programming help-seeking.  \nAbstract  \nAs students increasingly use LLM tutors in computer science education, one question becomes especially important: what kind of response helps a student continue productively? Prior work has studied how students use LLMs in computer science education, but less is known about how tutoring response styles are associated with student follow-up across programming help-seeking contexts. This paper analyzes StudyChat (UMass, 2026), a public dataset of student and ChatGPT tutoring conversations from an artificial intelligence course. We transformed StudyChat into 16,851 assistant-response interactions from 203 students and 2,214 conversations. Using local LLM-assisted annotation with Gemma 4, we labeled student help-seeking situations, student state, assistant response style, and student next-turn outcome. Human validation showed 82% agreement with the LLM-assisted labels (Cohen’s 􀁞 = . 74) . We analyzed productive continuation and unresolved continuation across the full dataset and across help-seeking contexts. Globally, response style was significantly associated with productive continuation,􀁪2 (7) = 100.39, 􀀿 \u003C .001, 􀀫 = .078, and unresolved continuation,􀁪2 (7) = 125.77, 􀀿 \u003C .001, 􀀫 = . 087, though effect sizes were small. Verification feedback had the highest productive-continuation rate (82.4%), while direct answers had the lowest (62.7%) . Descriptively,  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License. SIGCSE TS 2027, Sacramento, CA  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 978-1-4503-XXXX-X/2018/06  \n[https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \nresponse-style score ranges were smallest in low-confusion conceptual contexts (.017) and largest in high-cognitive-load contexts (.203) . More detailed comparisons showed situation-dependent response patterns. For example, stepwise guidance was followed by greater confusion decrease in high-cognitive-load code requests, while direct answers were followed by more unresolved continuation in high-load debugging. These findings support context-aware evaluation and design of AI tutoring responses for programming education.  \nCCS Concepts  \n• Social and professional topics → Computer science education; • Applied computing → Computer-assisted instruction.  \nKeywords  \nLLM Tutoring, Computer Science Education, Programming Education, Student Help-Seeking, Response Strategies, Conversational Analysis, Semantic Annotation, AI Education, Human-AI Interaction, Learning Analytics  \nACM Reference Format:  \nMohammad Fahim Abrar, Shyala Sharmin, and Roghayeh Leila Barmaki.  \n2027. When LLM Tutoring Responses Work: Evidence from Student Programming Conversations. In Proceedings of The Technical Symposium on Computer Science Education (SIGCSE TS 2027). ACM, New York, NY, USA, 7 pages. [https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \n1 Introduction  \nComputer Science students increasingly turn to large language models (LLMs) when they are stuck, confused, or trying to move  \nforward in programming tasks. In these moments, the response style of the assistant matter","cbCaihLogOWnUxWb","https://ap.wps.com/l/cbCaihLogOWnUxWb","pdf",1323691,2,1,7,"English","en",105,"# Abstract\n# Introduction\n# Background and Related Work\n# Dataset and Annotation\n# Analysis of Response Styles\n# Results: Productive vs Unresolved Continuation\n# Context-Dependent Findings\n# Implications for AI Tutoring Design","[{\"question\":\"What dataset and how many interactions are analyzed to study tutoring response effectiveness?\",\"answer\":\"The study uses the StudyChat dataset of real student–ChatGPT tutoring conversations from a university AI course. It transforms the dataset into 16,851 assistant-response interactions from 203 students and 2,214 conversations.\"},{\"question\":\"How are student help-seeking situations and outcomes labeled in the study?\",\"answer\":\"Student help-seeking situations, student state, assistant response style, and next-turn outcomes are labeled using local LLM-assisted annotation with Gemma 4. Human validation reports 82% agreement with the labels.\"},{\"question\":\"Which tutoring response style shows the highest productive continuation rate, and which is lowest?\",\"answer\":\"Verification feedback has the highest productive-continuation rate (82.4%), while direct answers have the lowest (62.7%). Effects are statistically significant but with small effect sizes.\"}]",1784201438,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-llm-tutoring-responses-work-evidence-from-student-programming-conversations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-llm-tutoring-responses-work-evidence-from-student-programming-conversations/85155/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What dataset and how many interactions are analyzed to study tutoring response effectiveness?","Question",{"text":75,"@type":76},"The study uses the StudyChat dataset of real student–ChatGPT tutoring conversations from a university AI course. It transforms the dataset into 16,851 assistant-response interactions from 203 students and 2,214 conversations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are student help-seeking situations and outcomes labeled in the study?",{"text":80,"@type":76},"Student help-seeking situations, student state, assistant response style, and next-turn outcomes are labeled using local LLM-assisted annotation with Gemma 4. Human validation reports 82% agreement with the labels.",{"name":82,"@type":73,"acceptedAnswer":83},"Which tutoring response style shows the highest productive continuation rate, and which is lowest?",{"text":84,"@type":76},"Verification feedback has the highest productive-continuation rate (82.4%), while direct answers have the lowest (62.7%). Effects are statistically significant but with small effect sizes.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]