[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85607-en":3,"doc-seo-85607-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85607,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Question Type, Cognitive Load, and CEFR Alignment Evaluating LLM-Generated EFL Grammar Drill Exercises","This study evaluates the pedagogical viability of LLM-generated English as a Foreign Language (EFL) learning content. Using log data from Japanese junior high school students practicing on an LLM-driven grammar drilling application, the research analyzes how question modalities affect performance and whether theoretical localised CEFR difficulty tiers match empirical task difficulty. Results show a clear cognitive-load hierarchy: multiple-choice is lowest, cloze tasks hinder active recall most, and drag-and-drop causes the largest time penalties.","Question Type, Cognitive Load, and CEFR Alignment: Evaluating LLM-Generated EFL Grammar Drill Exercises  \nSteve WOOLLASTONa*, Brendan FLANAGANb, Yuko TOYOKAWAa, & Hiroaki OGATAa  \naKyoto University, Japan  \nbRitsumeikan University, Japan  \n*[s.m.woollaston@gmail.com](s.m.woollaston@gmail.com)  \nAbstract: This study evaluates the pedagogical viability of LLM-generated English asa Foreign Language ( EFL) learning content. Utilising log data from Japanese junior high school students practicing on a grammar drilling application, we analysed how different question modalities impact student performance and whether theoretical localised CEFR difficulty tiers accurately predict empirical task difficulty. Results reveal a clear performance hierarchy: multiple-choice questions carried the lowest cognitive load, cloze tasks posed the greatest barrier to active recall, and drag-and-drop exercises incurred the heaviest time penalties. Furthermore, learner data validated the CEFR-J grammar framework, showing a steady decline inaccuracy and increased response times as proficiency levels advanced. These findings demonstrate that LLMs can successfully generate learning content, while highlighting the need for developers to strategically sequence question modalities to transition learners from passive recognition to active linguistic production.  \nKeywords: Grammar drill, EFL, log data, second language learning, CALL CEFR-J  \n1. Introduction  \nSecond Language Acquisition (SLA) is being transformed by Computer-Assisted Language Learning (CALL) systems, which offer learners the benefit of unlimited, self-directed, and self-paced practice sessions anywhere at any time ( Levy & Stockwell, 2013) . However, developing these digital learning systems so they are effective and aligned to established frameworks can be difficult for educators and developers. Handcrafting learning content demands significant pedagogical expertise, making content creation both time-consuming and expensive. Consequently, traditional CALL software often relies on static, limited item banks; when students exhaust these resources, they face repetitive exercises that lead to boredom, diminished learning outcomes and motivation.  \nTo overcome these scalability limitations, Large Language Models ( LLMs) are increasingly used to automate content generation ( Park & Derakhshan, 2026) . LLMs can instantly generate vast, linguistically diverse, personalised, and contextually rich content. However, the rapid deployment of AI-generated content raises pedagogical concerns regarding proficiency alignment (Zhang & Huang, 2024) . It remains unclear whether LLM-generated material reliably matches theoretical proficiency frameworks, or if it inadvertently introduces unintended friction for language learners. Furthermore, additional empirical research is required to evaluate how this generated content performs within localised English as a Foreign Language ( EFL) contexts such as Japanese secondary education. Popular CALL apps (e.g. , Duolingo) rely heavily on closed-ended question formats (e.g. , multiple-choice ( MCQ) , drag-and-drop, and cloze tasks) to provide immediate feedback, yet few studies have compared the direct performative impacts of these distinct question modalities on actual students (Syahid, 2018) .  \nTo address these gaps, this study utilises Grammar Gretel, an EFL grammar drilling application driven by a dataset of over 55 ,000 LLM generated English-Japanese sentence pairs mapped to the CEFR-J framework. By analysing the log data of Japanese junior highschool students, we evaluate the real-world viability of LLM-generated grammar drill practice. Specifically, this study investigates the following research questions:  \n1. How do different question modalities ( MCQ, drag-and-drop, cloze) impact student accuracy and response times?  \n2. Do generated CEFR sentence levels accurately predict empirical task difficulty for learners?  \n2. Related Work  \nWithin SLA research, the deba","cbCaiv6RnXLFz615","https://ap.wps.com/l/cbCaiv6RnXLFz615","pdf",610398,3,1,10,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"How do different question modalities affect student performance in the study?\",\"answer\":\"Multiple-choice questions produce the lowest cognitive load, cloze tasks create the greatest barrier to active recall, and drag-and-drop exercises lead to the heaviest time penalties.\"},{\"question\":\"Does the CEFR-J framework align with the real task difficulty observed in learner data?\",\"answer\":\"Yes. Learner data supports the CEFR-J grammar framework, showing decreasing inaccuracy and increasing response times as proficiency levels advance.\"},{\"question\":\"What guidance does the study provide for using LLMs to generate grammar drill exercises?\",\"answer\":\"LLMs can generate viable learning content, but developers should sequence question modalities strategically to move learners from passive recognition toward active linguistic production.\"}]",1784204894,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"question-type-cognitive-load-and-cefr-alignment-evaluating-llm-generated-efl-grammar-drill-exercises","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/question-type-cognitive-load-and-cefr-alignment-evaluating-llm-generated-efl-grammar-drill-exercises/85607/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How do different question modalities affect student performance in the study?","Question",{"text":75,"@type":76},"Multiple-choice questions produce the lowest cognitive load, cloze tasks create the greatest barrier to active recall, and drag-and-drop exercises lead to the heaviest time penalties.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Does the CEFR-J framework align with the real task difficulty observed in learner data?",{"text":80,"@type":76},"Yes. Learner data supports the CEFR-J grammar framework, showing decreasing inaccuracy and increasing response times as proficiency levels advance.",{"name":82,"@type":73,"acceptedAnswer":83},"What guidance does the study provide for using LLMs to generate grammar drill exercises?",{"text":84,"@type":76},"LLMs can generate viable learning content, but developers should sequence question modalities strategically to move learners from passive recognition toward active linguistic production.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]