[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160535-en":3,"doc-seo-160535-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160535,962085320529,"Sarah ","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Responsible AI for Test Equity and Quality - The Duolingo English Test as a Case Study","Artificial intelligence (AI) enables assessment efficiencies such as automated item generation and scoring of spoken and written responses, while also introducing risks including bias in AI-generated content. Responsible AI (RAI) practices mitigate these risks by guiding how evidence is gathered for valid test score interpretations. This chapter links RAI to test quality and test equity, using the Duolingo English Test (DET) as an AI-powered high-stakes case study. It details DET RAI standards and examples aligned with validity, reliability, fairness, privacy and security, and transparency and accountability.","Responsible AI for Test Equity and Quality:  \nThe Duolingo English Test as a Case Study  \nJill Bursteina, 1, Geoffrey T. LaFlaira, Kevin Yanceya, and Alina A. von Daviera  \nDuolingo, Inc.  \n5900 Penn Avenue  \nPittsburgh, PA 15206 {jill, geoff, kevin, [avondavier}@duolingo.com](avondavier}@duolingo.com)  \nRavit Dotanb,2  \nTechBetter LLC  \n[ravit@techbetter.ai](ravit@techbetter.ai)  \nChapter under review for the Handbookfor Assessment in the Service of Learning (Editorial Team: Eleanor Armour-Thomas, Eva L. Baker, Howard Everson, Edmund W. Gordon,  \nSteve Sireci, and Eric Tucker)  \n1 Corresponding author.  \n2 [https://www.ravitdotan.com/](https://www.ravitdotan.com/)  \nAbstract  \nArtificial intelligence (AI) creates opportunities for assessments, such as efficiencies for item generation and scoring of spoken and written responses. At the same time, it poses risks (such as bias in AI-generated item content) . Responsible AI (RAI) practices aim to mitigate risks associated with AI. This chapter addresses the critical role of RAI practices in achieving test quality (appropriateness oftest score inferences), and test equity (fairness to all test takers)—key principles in this volume. To illustrate, the chapter presents a case study using the Duolingo English Test (DET)–an AI-powered, high-stakes English language assessment. The chapter discusses the DET RAI standards, their development and their relationship to domain-agnostic RAI principles. Further, it provides examples of specific RAI practices, showing how these practices meaningfully address the ethical principles of validity and reliability, fairness, privacy and security, and transparency and accountability standards to ensure test equity and quality.  \n1. Introduction  \nTest quality is achieved through evidence gathering that confirms an assessment's suitability for its intended purpose. Test equity is attained when test scores are fair – specifically, they do not favor or disadvantage a particular group. Classical argument-based test validity theory supports test quality and equity through a chain of inferences. Inferences assume mechanisms (e.g., relevant task types) through which evidence collection ensures an appropriate test score interpretation (Chapelle et al, 2008; Kane, 1992) . These include: domain definition (task types represent the target domain as defined), evaluation (test scores reflect language ability), generalization (test scores are reliable), explanation (test scores are attributable to the construct), extrapolation (test score are related to other language criteria), utilization (test scores are interpretable and meaningful for their purpose) . To maintain the validity of AI-powered assessments, it is essential to evaluate AI capabilities, as they impact evidence collection, measurement, and, ultimately, test quality and equity.  \nWith recent generative artificial intelligence (AI) advances, AI-powered assessments are becoming increasingly common. AI for assessment offers many advantages, such as automated scoring of writing and speaking, and creating larger item banks through efficiencies in automated item generation. However, there are risks. For example, bias in AI-generated item content can impact test-taker outcomes (e.g., Belzak et al., 2023; Johnson et al, 2022); this can, potentially, lead to test inequity and diminished test quality. Therefore, AI-powered assessment calls for alignment with human-centered AI values which are enacted through responsible AI guidelines and standards (von Davier A. & Burstein, to appear; Burstein, 2023; Auernhammer, 2020; The IEEE Global Initiative on Ethics of Autonomous and Intelligent Systems, 2017) . Though some risks to validity are similar across traditional and AI-based assessments, AI introduces unique  \nrisks. While assessment standards address AI to some extent (American Educational Research Association (AERA), American Psychological Association (APA), & National Council on Measurement in Education (NCME)","cbCaiv7Jar2dhVgL","https://ap.wps.com/l/cbCaiv7Jar2dhVgL","pdf",1041072,5,1,45,"English","en",105,"# Abstract\n# Introduction\n# Background and Related Work","[{\"question\":\"How do responsible AI practices support test quality in AI-powered assessments?\",\"answer\":\"RAI helps ensure evidence is gathered to support an appropriate chain of inferences for test score interpretation, strengthening validity arguments. It also emphasizes evaluating AI capabilities that affect measurement and evidence collection.\"},{\"question\":\"What is the relationship between test equity and AI bias in assessment?\",\"answer\":\"AI-generated item bias can disadvantage specific groups, creating test inequity. The chapter addresses fairness through RAI standards and practices that reduce bias risks.\"},{\"question\":\"What does the case study on the Duolingo English Test (DET) cover?\",\"answer\":\"The chapter presents DET RAI standards, explains how they were developed, and illustrates how they are validated and applied to support test equity and quality. It also discusses implications and known limitations.\"}]","Responsible AI for Test Equity and Quality - The Duolingo English Test as a Case Study | PDF",1788068599,113,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"responsible-ai-for-test-equity-and-quality-the-duolingo-english-test-as-a-case-study","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/responsible-ai-for-test-equity-and-quality-the-duolingo-english-test-as-a-case-study/160535/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-09-04","2026-08-30",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"How do responsible AI practices support test quality in AI-powered assessments?","Question",{"text":77,"@type":78},"RAI helps ensure evidence is gathered to support an appropriate chain of inferences for test score interpretation, strengthening validity arguments. It also emphasizes evaluating AI capabilities that affect measurement and evidence collection.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What is the relationship between test equity and AI bias in assessment?",{"text":82,"@type":78},"AI-generated item bias can disadvantage specific groups, creating test inequity. The chapter addresses fairness through RAI standards and practices that reduce bias risks.",{"name":84,"@type":75,"acceptedAnswer":85},"What does the case study on the Duolingo English Test (DET) cover?",{"text":86,"@type":78},"The chapter presents DET RAI standards, explains how they were developed, and illustrates how they are validated and applied to support test equity and quality. It also discusses implications and known limitations.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]