[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82651-en":3,"doc-seo-82651-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82651,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","AIriskEval-edu-db2: 新的K–12教育解释风险评估数据集","AIriskEval-edu-db2 introduces a new dataset for training and evaluating auditors based on LLMs to perform explainable pedagogical risk assessment of K–12 instructional explanations. The dataset contains 1,639 explanations generated from 170 curated ScienceQA questions spanning science, language arts, and social sciences. Each question pairs a human teacher explanation with 11 LLM-simulated teacher-profile explanations tied to distinct pedagogical risks. A comprehensive rubric covers factual precision, depth/completeness, focus/relevance, student-level appropriateness, and ideological bias, with structured explainability annotations validated by experts. Validation compares proprietary frontier models against a local Llama 3.1 8B for risk detection and explainability assessment, examining privacy-preserving deployability via fine-tuning.","AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations  \nJavier Irigoyen∗ , Roberto Daza∗†, Francisco Jurado†, Julian Fierrez∗ ,  \nRuben Tolosana∗ , Alvaro Ortigosa†, Enrique Blas∗ , and Aythami Morales∗ ,‡  \n∗ BiometricsAI, Universidad Auto´noma de Madrid (UAM), Spain  \n†GHIA, Universidad Auto´noma de Madrid (UAM), Spain  \n‡ Universidad de Las Palmas de Gran Canaria (ULPGC), Spain  \nCorresponding author: roberto.daza@uam.es  \narXiv :2607 .0 1934v 1 [ cs .CL] 2 Jul 2026  \nAbstract—This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk assessment in instructional content for grades K–12. The dataset comprises 1,639 explanations from 170 curated ScienceQA questions, covering science, language arts, and social sciences. For each question, the dataset includes an explanation written by a human teacher alongside 11 explanations generated by LLM-simulated teacher profiles associated with distinct pedagogical risks. We propose a comprehensive risk rubric aligned with established educational standards that covers five complementary dimensions: factual precision, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. A key contribution is the addition of 785 explanations with structured explainability annotations, including risk localization and risk description. The annotations are produced through a semi-automatic process with expert teacher validation. Finally, we present validation experiments comparing state-of-the-art proprietary models with a lightweight local Llama 3.1 8B model in both the pedagogical risk detection and the explainability assessment. These experiments evaluate whether supervised fine-tuning on AIriskEval-edu-db2 enables a locally deployable model to approach or outperform stronger frontier models while preserving privacy in educational auditing and assessment tasks.  \nI. INTRODUCTION  \nLarge Language Models (LLMs) are increasingly being deployed in human-facing systems where raw capability must be complemented by oversight, automatic monitoring, and risk-aware quality assurance. In educational technology, these models can answer K–12 questions with high precision [1], which has raised growing interest in their use as both tutors and as automatic evaluators of instructional explanations produced by humans or AI. At the same time, because such systems can produce unsafe, unreliable, or misleading outputs, their deployment also raises concerns that are closely related to the themes of automatic monitoring, human factors, artificial intelligence, and the societal impact of security technology [2] . In this sense, educational AI can also be viewed as a securitytechnology problem: it requires monitoring mechanisms able  \nThis research was supported by Ctedra ENIA UAM-VERIDASen IA Responsable (NextGenerationEU PRTR TSI-100927-2023-2), M2RAI (PID2024-160053OB-I00, MICIU/FEDER), TRUST-ID (PID2025- 173396OB-I00, MICIU/AEI and the EU) and PowerAI+ (SI4/PJI/2024-00062, Comunidad de Madrid and UAM) . Javier Irigoyen is supported by an FPI fellowship from MINECO/FEDER.  \nto detect harmful, untrustworthy, or otherwise risky language outputs before or during deployment.  \nRecent studies show that LLMs, especially when taskadapted, can approximate key tutoring behaviors [3], while industrial efforts such as LearnLM [4] and studies on the adaptation of instructional roles [5] further support this trend. At the same time, LLMs introduce well-documented risks [6], which are especially consequential in K–12 settings. This motivates automatic pedagogical assessors capable of detecting pedagogical and epistemic risks under established educational frameworks [7] .  \nPrevious work suggests that fine-tuning LLMs in rubricbased educational datasets can improve the reliability of theevaluator [8], [9] . However, current public resources remain limited: few target evaluat","cbCaimhV9Rp2k5Qz","https://ap.wps.com/l/cbCaimhV9Rp2k5Qz","pdf",524323,2,1,6,"English","en",105,"# Introduction\n## Motivation and Problem Setting\n## Related Work and Limitations\n## Contributions and Extensions\n## Dataset Construction and Annotation","[{\"question\":\"What is AIriskEval-edu-db2 designed for?\",\"answer\":\"AIriskEval-edu-db2 is designed to train and evaluate LLM-based auditors for explainable pedagogical risk assessment in K–12 instructional explanations.\"},{\"question\":\"How is the dataset structured around teacher explanations and pedagogical risks?\",\"answer\":\"For each of the 170 curated ScienceQA questions, the dataset includes a human teacher explanation plus 11 explanations generated by LLM-simulated teacher profiles, each associated with distinct pedagogical risks.\"},{\"question\":\"What rubric dimensions are used to assess pedagogical risks?\",\"answer\":\"The proposed rubric covers five dimensions: factual precision, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias.\"}]",1784182069,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"airiskeval-edu-db2-new-dataset-for-risk-assessment-in-ai-mediated-k12-educational-explanations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/airiskeval-edu-db2-new-dataset-for-risk-assessment-in-ai-mediated-k12-educational-explanations/82651/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is AIriskEval-edu-db2 designed for?","Question",{"text":75,"@type":76},"AIriskEval-edu-db2 is designed to train and evaluate LLM-based auditors for explainable pedagogical risk assessment in K–12 instructional explanations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the dataset structured around teacher explanations and pedagogical risks?",{"text":80,"@type":76},"For each of the 170 curated ScienceQA questions, the dataset includes a human teacher explanation plus 11 explanations generated by LLM-simulated teacher profiles, each associated with distinct pedagogical risks.",{"name":82,"@type":73,"acceptedAnswer":83},"What rubric dimensions are used to assess pedagogical risks?",{"text":84,"@type":76},"The proposed rubric covers five dimensions: factual precision, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]