[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85235-en":3,"doc-seo-85235-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85235,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Knowledge Distillation for Automated AI Tutor Evaluation","Rapid adoption of Large Language Models (LLMs) in K-12 and higher education has outpaced dependable evaluation methods for pedagogical quality. To support automated assessment of AI tutors, the work proposes FATE (FLC AI Tutor Evaluator), an 8B-parameter model aligned to BEA 2025’s four evaluation tracks: Mistake Identification, Mistake Location, Guidance, and Actionability. With limited labeled data, knowledge distillation from a frontier LLM adds supervision, delivering absolute gains up to 22.63 percentage points. FATE is further benchmarked on responses from major commercial models.","Knowledge Distillation for Automated AI Tutor Evaluation  \nTahmid Al Hannan, Diego Garcia, Alex Njoroge, Suha Al Juboori, Tarek Sakakini  \nFolsom Lake College  \nFolsom, CA, USA  \n{w2124644, w2079995, [w2140050}@apps.losrios.edu](w2140050}@apps.losrios.edu), {aljubos, [sakakit}@flc.losrios.edu](sakakit}@flc.losrios.edu)  \narXiv :2607 . 10647v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nThe rapid integration of Large Language Models (LLMs) into K-12 and higher education has outpaced the development of reliable methods for evaluating their pedagogical quality. As the research community starts to explore the space of automating evaluation of AI tutors, we introduce FATE (FLC AI Tutor Evaluator), a specialized 8B-parameter language model designed to evaluate AI tutors. Aligned with the four core evaluation tracks from the BEA 2025 Shared Task, our model assesses pedagogical ability across Mistake Identification, Mistake Location, Guidance, and Actionability. Because pedagogical evaluation is a specialized task with limited labeled data, we leverage knowledge distillation from a frontier LLM to generate additional supervision, yielding absolute performance gains up to 22.63 percentage points.  \nFinally, we demonstrate FATE’s utility as an automated evaluator by benchmarking instructional responses generated by popular commercial models, including ChatGPT, Claude, Gemini, and DeepSeek. On average, we have found that Gemini 2.5 Flash perfomed best (82.88%), then ChatGPT 5.5 Instant (80.75%), followed by DeepSeek V4 Flash (80.13%) and Claude Sonnet 4 .6 (74 .00%) .  \n1 Introduction  \nLarge language models (LLMs) are becoming increasingly prevalent in education, with AI usage in observed K–12 classrooms increasing by 84% from 2024 to 2025 (McGehee) . As AI tutors become more widely adopted, reliable methods for evaluating their pedagogical quality are increasingly important. The BEA 2025 Shared Task introduced a benchmark for evaluating AI tutors across multiple pedagogical dimensions (Figure 1) (Kochmar et al., 2025a) . However, human evaluation is costly and difficult to scale, while automatic metrics often fail to capture qualities such as guidance, actionability, and instructional effectiveness (MarquezCarpintero et al., 2025) . This motivates automated  \nConversation History:  \nTutor : 1 is the first smallest digit in the number 3591. Which is the next smallest digit in 3591 other than 1?  \nStudent: 9 because it is a tens place and 3 is a thousands place and 5 is a hundreds place .  \nTutor Response to Final Student Utterance:  \nGood Response:  \nI see…let's remember that when comparing digits, we need to think about their actual value, not the place value, so which digit is smaller, 3, 5, or 9, if we just look at them as single numbers?  \nMI: Yes ML: Yes PG: Yes Actionability: Yes  \nOkay Response:  \nI see…but remember we're focusing on the value of single digits, not their place in the number . So, the next smallest digit after 1 in the number 3591 is actually 3 .  \nMI: Yes ML: Yes PG: Yes Actionability: No  \nBad Response:  \nThere is a misunderstanding with place values because when we look at the number 3591, the digit 9 is actually in the hundreds place, not because it's 10, but because it comes after the digit 5 in the hundreds place .  \nMI: Yes ML: Somewhat PG: Somewhat Actionability: No  \nFigure 1: Example dataset point from BEA 2025 . Below each AIT response is the human evaluation using four metrics: Mistake Identification (MI), Mistake Location (ML), Providing Guidance (PG), and Actionability.  \nevaluators capable of assessing tutoring interactions at scale (Tack et al., 2023) .  \nRecent evaluation frameworks extend beyond traditional benchmarks. Multimodal environments such as TutorBench provide comprehensive evaluations but require substantial computational resources and are often mismatched with the textbased chatbot interfaces most widely used today (Srinivasa et al., 2025 ; Kajan et al., 2025) . Textbased tutoring has also bee","cbCaiswhIiOy0vD2","https://ap.wps.com/l/cbCaiswhIiOy0vD2","pdf",206031,1,5,"English","en",105,"# Abstract\n# Introduction\n# Prior Work","[{\"question\":\"What is FATE and what does it evaluate in AI tutors?\",\"answer\":\"FATE (FLC AI Tutor Evaluator) is a specialized 8B-parameter language model for evaluating AI tutors. It assesses four pedagogical dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability.\"},{\"question\":\"How does the approach improve performance when labeled data is limited?\",\"answer\":\"The method uses knowledge distillation from a frontier LLM to generate additional supervision. This enables training FATE under limited labeled data and yields absolute performance gains up to 22.63 percentage points.\"},{\"question\":\"How is FATE used to evaluate commercial AI tutors?\",\"answer\":\"FATE benchmarks instructional responses produced by commercial models such as ChatGPT, Claude, Gemini, and DeepSeek. The document reports that Gemini 2.5 Flash achieved the highest average performance (82.88%) among the compared systems.\"}]",1784201907,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"knowledge-distillation-for-automated-ai-tutor-evaluation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/knowledge-distillation-for-automated-ai-tutor-evaluation/85235/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is FATE and what does it evaluate in AI tutors?","Question",{"text":75,"@type":76},"FATE (FLC AI Tutor Evaluator) is a specialized 8B-parameter language model for evaluating AI tutors. It assesses four pedagogical dimensions: Mistake Identification, Mistake Location, Providing Guidance, and Actionability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the approach improve performance when labeled data is limited?",{"text":80,"@type":76},"The method uses knowledge distillation from a frontier LLM to generate additional supervision. This enables training FATE under limited labeled data and yields absolute performance gains up to 22.63 percentage points.",{"name":82,"@type":73,"acceptedAnswer":83},"How is FATE used to evaluate commercial AI tutors?",{"text":84,"@type":76},"FATE benchmarks instructional responses produced by commercial models such as ChatGPT, Claude, Gemini, and DeepSeek. The document reports that Gemini 2.5 Flash achieved the highest average performance (82.88%) among the compared systems.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]