[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82963-en":3,"doc-seo-82963-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82963,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming","Large language models are increasingly considered as AI tutors, but K–12 deployment raises privacy, cost, and dependence on proprietary systems. Small language models (SLMs) offer a practical alternative, yet choosing an appropriate model remains difficult when the target domain is missing from training data. CSTutorBench benchmarks language models as computer science tutors in VEX VR, using 17 scenario questions, an 8-criterion rubric, and an LLM-as-judge evaluation pipeline. Results show strong surface quality but weaker deeper pedagogical behaviors, with prompt revision improving most models.","CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming  \nH. Chad Lane1 , Bryson Kageler1  \n1 University of Illinois Urbana-Champaign, Champaign, IL, USA  \nAbstract  \nLarge language models are increasingly explored as AI tutors, yet deploying them in K–12 settings raises concerns around privacy, cost, and reliance on proprietary models. Small language models (SLMs) offer a promising alternative, but selecting the right model for a specific educational context remains difficult, particularly when the target domain, such as block-based programming, is largely absent from model training data. We introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics environment. The benchmark comprises 17 scenario-based questions scored against a pedagogical rubric grounded in established tutoring and feedback research, with a human-in-the-loop LLM-as-judge pipeline for evaluation. Preliminary findings across 11 models (4B–120B parameters) reveal that models perform well on surface-level criteria such as vocabulary and tone but struggle with deeper pedagogical behaviors, particularly avoiding answer leakage and engaging with student debugging histories. In our sample, model family and instruction-tuning approach appear to be better predictors of tutoring quality than parameter count alone, though the small number of models limits the strength of this conclusion. A targeted prompt revision grounded in recent educational prompt engineering research improved scores for 10 of 11 models. These results underscore the value of context-specific, pedagogically grounded benchmarks for SLM selection in educational deployment.  \nKeywords  \nsmall language models, benchmark, intelligent tutoring, block-based programming, LLM-as-judge  \n1. Introduction  \nLarge language models (LLMs) have demonstrated considerable potential as educational tools, particularly for tasks that involve generating explanations, providing feedback, and engaging learners in dialogue [1] . For computer science education, it is particularly appealing that many models are trained on large code corpora, meaning they can reliably produce syntactically correct solutions, trace through program logic, and explain programming concepts with impressive fluency [2] . But, echoing a classic lesson learned in AI in Education showing that it is unwise to simply wrap a tutor around an expert system [3]: we should not expect an LLM that is good at coding to be good at teaching people how to code. Effective tutoring involves a delicate balance of providing guidance, but allowing learners to maintain a sense of agency [4, 5, 6] . Although some benchmarks do focus on the educational capabilities of LLMs [7] and others on tutoring [8], there continues to be a need to more directly assess the tutoring and coaching capabilities of modern LLMs and investigate new methods for customizing their behavior.  \nThese challenges are amplified when the target population consists of young learners working in block-based programming environments such as Scratch, VEX, or Snap!. Block-based environments are designed specifically for education and are largely absent from the training data of even the most capable models. At the same time, practical constraints in K-12 settings around privacy, limited budgets, and the need for a high degree of local control, all point toward small language models (SLMs) asa viable path for addressing these needs. But SLMs vary widely in their behavior and there is little guidance to deciding which 8B or 30B model to deploy as a classroom tutor. In this paper, we introduce CSTutorBench, a benchmark for evaluating language models as CS tutors in VEX VR, a block-based robotics simulation for middle school students. The benchmark includes 17 scenario-based questions, an 8-criterion pedagogical rubric, and an automated LLM-as-judge evaluation pipeline. We present preliminary results comparing models from 4B ","cbCaigmNK46zBE2K","https://ap.wps.com/l/cbCaigmNK46zBE2K","pdf",1943168,4,1,9,"English","en",105,"# 1. Introduction\n# 2. CSTutorBench\n## Task Domain\n## Dataset","[{\"question\":\"Why is it challenging to use small language models as tutors for block-based programming in K–12 settings?\",\"answer\":\"Block-based platforms such as VEX VR are designed for education and are largely absent from common training data. This makes it difficult to know which SLM will tutor effectively despite constraints like privacy, limited budgets, and the need for local control.\"},{\"question\":\"What is CSTutorBench and how is it evaluated?\",\"answer\":\"CSTutorBench is a benchmark for evaluating language models as CS tutors in VEX VR. It uses 17 scenario-based questions scored with a pedagogical rubric and applies an LLM-as-judge pipeline to automate evaluation.\"},{\"question\":\"What do preliminary results indicate about tutoring quality across different SLMs?\",\"answer\":\"Models tend to perform well on surface-level criteria such as vocabulary and tone, but struggle with deeper pedagogical behaviors like avoiding answer leakage and engaging with students’ debugging histories. The paper also reports that instruction-tuning and model family may predict tutoring quality better than parameter count alone.\"}]",1784184358,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cstutorbench-benchmarking-small-language-models-as-tutors-for-block-based-programming","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/cstutorbench-benchmarking-small-language-models-as-tutors-for-block-based-programming/82963/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is it challenging to use small language models as tutors for block-based programming in K–12 settings?","Question",{"text":75,"@type":76},"Block-based platforms such as VEX VR are designed for education and are largely absent from common training data. This makes it difficult to know which SLM will tutor effectively despite constraints like privacy, limited budgets, and the need for local control.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is CSTutorBench and how is it evaluated?",{"text":80,"@type":76},"CSTutorBench is a benchmark for evaluating language models as CS tutors in VEX VR. It uses 17 scenario-based questions scored with a pedagogical rubric and applies an LLM-as-judge pipeline to automate evaluation.",{"name":82,"@type":73,"acceptedAnswer":83},"What do preliminary results indicate about tutoring quality across different SLMs?",{"text":84,"@type":76},"Models tend to perform well on surface-level criteria such as vocabulary and tone, but struggle with deeper pedagogical behaviors like avoiding answer leakage and engaging with students’ debugging histories. The paper also reports that instruction-tuning and model family may predict tutoring quality better than parameter count alone.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]