[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86399-en":3,"doc-seo-86399-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86399,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Learning in Blocks A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning","Most digital language learning relies on discrete-item quizzes that measure recall rather than real conversational ability, causing learners to progress despite persistent gaps in grammar and vocabulary use during interaction. Learning in Blocks introduces a CEFR-rubric grounded framework where heterogeneous multi-agent debate produces reliable conversation scoring and targeted recommendations. Two stages—scoring via role-specialized debate and judge consensus, then review planning using weaknesses—combine 70% mastery progression with spaced review. Benchmarks on CEFR A2 conversations and an 8-week 180-learner study show improved outcomes over feedback alone.","arXiv :2604 .22770v2 [ cs .CY] 13 Jul 2026  \nLearning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning  \nNicy Scaria 1⋆ , Silvester John Joseph Kennedy 1 ,2 ,3⋆ , and Deepak  \nSubramani 1  \n1 Indian Institute of Science, India  \n2 Talking Yak, India  \n3 Indian Institute of Technology Patna, India  \n[nicyscaria@iisc.ac.in](nicyscaria@iisc.ac.in)  \nAbstract. Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progression is driven by quiz performance, learners can advance despite persistent gaps in using grammar and vocabulary during interaction. Recent work on LLM-based judging suggests a path toward scoring open-ended conversations, but using interaction evidence to drive progression and review requires scoring protocols that are reliable and validated. We introduce Learning in Blocks, a framework that grounds progression in demonstrated conversational competence evaluated using CEFR-aligned rubrics. The framework employs heterogeneous multiagent debate (HeteroMAD) in two stages: a scoring stage where rolespecialized agents independently evaluate Grammar, Vocabulary, and Interactive Communication, engage in debate to address conflicting judgments, and a judge synthesizes consensus scores; and a recommendation stage that identifies specific grammar skills and vocabulary topics for targeted review. Progression requires demonstrating 70% mastery, and spaced review targets identified weaknesses to counter skill decay. We benchmark four scoring and recommendation methods on CEFR A2 conversations annotated by ESL experts. HeteroMAD achieves a superior score agreement with a 0.23 degree of variation and recommendation acceptability of 90 .91% . An 8-week study with 180 CEFR A2 learners demonstrates that combining rubric-aligned scoring and recommendation with spaced review and mastery-based progression produces better learning outcomes than feedback alone.  \nKeywords: Multi-Agent Debate · Personalized Adaptive Learning · Mastery Based Learning.  \n⋆ These authors contributed equally.  \nThis is a preprint. The Version of Record of this contribution will be published in International Conference on Artificial Intelligence in Education (LNAI), and will be available shortly.  \n2 N. Scaria et al.  \n1 Introduction  \nDigital language learning curricula typically rely on discrete-item quizzes that test recall of grammar rules or vocabulary definitions, creating a disconnect between assessment formats (isolated recall tasks) and and what constitutes communicative proficiency (applied performance in context) . When progression is driven by quiz performance, learners can advance through content while gapsin applying Grammar, Vocabulary, and Interactive Communication persist, producing a Swiss cheese learning [17] pattern. Large language models have enabled conversational AI applications for language learning [16], allowing learners to engage in open-ended dialogue practice. However, using open-ended spoken interaction for assessment and progression requires scoring protocols that are reliable and validated against expert judgment. In the absence of such validation, platforms often default to quiz-based signals that are straightforward to standardize and score at scale, even when those signals do not capture whether learners can apply skills during interaction. Recent work on rubric-based LLM judging [31, 19] suggests a path toward scoring open-ended interaction more directly, creating an opportunity to ground progression decisions in demonstrated conversational competence rather than quiz performance.  \nThis challenge is especially acute for conversational proficiency, which requires coordinated control of grammatical accuracy, lexical appropriateness, and interactional success across multiple turns [13] . In language learning contexts, CEFR provides widely adopted descriptors that operationalize these di","cbCaigBcVBJy7hvd","https://ap.wps.com/l/cbCaigBcVBJy7hvd","pdf",943135,1,15,"English","en",105,"# Introduction\n## Assessment gaps in quiz-driven curricula\n## Rubric-based LLM judging and scoring reliability\n## Mastery-based progression and skill decay\n# Learning in Blocks Framework\n## CEFR-aligned scoring using HeteroMAD\n## Recommendation stage and targeted review\n## 70% mastery progression and spaced review\n# Experimental Evaluation","[{\"question\":\"Why do quiz-based language learning systems fail to reflect communicative proficiency?\",\"answer\":\"They test recall of isolated grammar or vocabulary items, so learners can advance even when they still cannot apply these skills in real conversation. This creates persistent performance gaps despite content progression.\"},{\"question\":\"How does Learning in Blocks evaluate conversational competence?\",\"answer\":\"It uses CEFR-aligned rubrics to score Grammar, Vocabulary, and Interactive Communication, produced through heterogeneous multi-agent debate. Role-specialized agents debate conflicting judgments, and a judge synthesizes consensus scores.\"},{\"question\":\"What determines progression and how does the framework prevent skill decay?\",\"answer\":\"Progression requires meeting a 70% mastery criterion on block-aligned evaluation. Spaced review targets identified weaknesses to re-elicitate and reinforce skills over time.\"}]",1784211495,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"learning-in-blocks-a-multi-agent-debate-assisted-personalized-adaptive-learning-framework-for-language-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/learning-in-blocks-a-multi-agent-debate-assisted-personalized-adaptive-learning-framework-for-language-learning/86399/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do quiz-based language learning systems fail to reflect communicative proficiency?","Question",{"text":75,"@type":76},"They test recall of isolated grammar or vocabulary items, so learners can advance even when they still cannot apply these skills in real conversation. This creates persistent performance gaps despite content progression.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Learning in Blocks evaluate conversational competence?",{"text":80,"@type":76},"It uses CEFR-aligned rubrics to score Grammar, Vocabulary, and Interactive Communication, produced through heterogeneous multi-agent debate. Role-specialized agents debate conflicting judgments, and a judge synthesizes consensus scores.",{"name":82,"@type":73,"acceptedAnswer":83},"What determines progression and how does the framework prevent skill decay?",{"text":84,"@type":76},"Progression requires meeting a 70% mastery criterion on block-aligned evaluation. Spaced review targets identified weaknesses to re-elicitate and reinforce skills over time.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]