[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82062-en":3,"doc-seo-82062-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82062,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","L2-Bench An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education","Despite fast adoption of AI in education, rigorous evaluation for AI-powered educational systems remains limited, especially for second language (L2) learning. L2-Bench introduces an open-source benchmark with 1,000+ task-response pairs to evaluate LLM performativity through learning-experience design principles, not only knowledge of them. It provides a validated taxonomy of 12 competencies and 31 subcompetencies, a rubric-based evaluation method, and a dataset producing reliable signals on strengths, weaknesses, and robustness across L2 scenarios, enabling more informed AIED use and governance.","L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second  \nLanguage Education  \nJames Edgell* 1 , Wm. Matthew Kennedy*2, 3 , Ben Knight 1 , Danielle Carvalho 1 , Martin Ku 1 , Isaac Pattis 1  \n1 Oxford University Press  \n2 Oxford Internet Institute, University of Oxford  \n3 King’s College London  \n[elt-bench@oup.com](elt-bench@oup.com)  \narXiv :2607 .08842v 1 [ cs .CY] 9 Jul 2026  \nAbstract  \nDespite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity = 4.42/5.00, criteria adequacy = 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks ∈(69.9%–73.4%) . L2-Bench provides education stakeholders better methods to make more informed decisions about realworld AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.  \n1 Introduction  \nDespite rapid adoption of AI systems in educational spaces (Digital Education Council 2024), very few evaluations for AI in educational (AIED) exist. Those that purport to cover this space often suffer from poor construct validity, are underpowered, or focus more on broad effects rather than instance-level performativity. The result is little short of a“wild, wild west”: a situation in which deployment is far outpacing evidence even of basic validation of AIED systems (Hudig, Kallina, and Singh 2026) .  \nThis AIED evaluations gap reflects the ongoing crisis in AI evaluations in general. However, it is more acute: although prior validation and efficacy studies of conventional  \n*These authors contributed equally.  \nCopyright © 2026, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \neducation technologies suggest that these predecessor systems have not meaningfully improved learning outcomes (Cuban 2001; Zawacki-Richter et al. 2019; Selwyn 2012, 2019), many AIED systems replicate or extend the same conventional system design patterns (Jurenka et al. 2024) . Such a headlong rush on the part of AIED developers and adopters without parallel development of novel evaluation methodologies creates the conditions for myriad interaction-, learning-group-, systemic-, and compounding harms (Bastani et al. 2024; Holmes and Miao 2023; Kasneci et al. 2023; Wachter, Mittelstadt, and Russell 2024; Holmes 2024; Kennedy and Campos 2025) . In the meantime, a generation of learners must endure laboratory conditions (Alcaras and Ricci 2025) in their pursuit of educational attainment.  \nDrawing inspiration from allied efforts to improve the state of the evaluations ecosystem (Reuel et al. 2024; Eriksson et al. 2025; Biderman et al. 2024; Weidinger et al. 2025; Schwartz et al. 2025; Bean et al. 2025; Paskov, Soder, and Smith 2025; Reuel et al. 2025), our work hastens progress in one particularly challenging subdomain of AI for education: second language (L2) educat","cbCainRtESV7pa3F","https://ap.wps.com/l/cbCainRtESV7pa3F","pdf",2718491,1,49,"English","en",105,"# Abstract\n# Introduction\n## Evaluation gap in AI for education\n## Motivation for L2-Bench\n## Benchmark contributions","[{\"question\":\"What problem does L2-Bench address in second language education and AI evaluation?\",\"answer\":\"L2-Bench targets the lack of rigorous, construct-valid evaluations for AI-powered educational (AIED) systems in second language (L2) education, where many existing evaluations are underpowered or focus on broad effects rather than task-level performativity.\"},{\"question\":\"How does L2-Bench evaluate LLMs differently from knowledge-only or outcome-only assessments?\",\"answer\":\"L2-Bench measures model performativity by checking how well models apply learning experience design principles for language learning and assessment, rather than merely testing knowledge of those principles or broad learning outcomes.\"},{\"question\":\"What are L2-Bench’s main contributions and what does the benchmark contain?\",\"answer\":\"L2-Bench contributes a validated taxonomy of 12 competencies and 31 subcompetencies, a rubric-based evaluation methodology, and a benchmark dataset of 1,000+ structured task-response pairs to support reliable comparisons across L2 education scenarios.\"}]",1784177950,123,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"l2-bench-an-evaluation-benchmark-for-measuring-llm-capabilities-in-second-language-education","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/l2-bench-an-evaluation-benchmark-for-measuring-llm-capabilities-in-second-language-education/82062/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does L2-Bench address in second language education and AI evaluation?","Question",{"text":75,"@type":76},"L2-Bench targets the lack of rigorous, construct-valid evaluations for AI-powered educational (AIED) systems in second language (L2) education, where many existing evaluations are underpowered or focus on broad effects rather than task-level performativity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does L2-Bench evaluate LLMs differently from knowledge-only or outcome-only assessments?",{"text":80,"@type":76},"L2-Bench measures model performativity by checking how well models apply learning experience design principles for language learning and assessment, rather than merely testing knowledge of those principles or broad learning outcomes.",{"name":82,"@type":73,"acceptedAnswer":83},"What are L2-Bench’s main contributions and what does the benchmark contain?",{"text":84,"@type":76},"L2-Bench contributes a validated taxonomy of 12 competencies and 31 subcompetencies, a rubric-based evaluation methodology, and a benchmark dataset of 1,000+ structured task-response pairs to support reliable comparisons across L2 education scenarios.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]