[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160530-en":3,"doc-seo-160530-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160530,687207022233,"Connor ","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",4,"Exam","Analysis of the Scoring and Reliability for the Duolingo English Test - Procedures and Reliability/Precision","Review of research supporting Duolingo English Test (DET) scoring methods and the reliability/precision of resulting scores. The document applies the Standards for Educational and Psychological Testing (AERA, APA, & NCME, 2014) together with DET documentation and analyzed questions on established scoring procedures, item generation and difficulty estimation, item selection for administration, and how item responses produce total scores and subscores. It further examines evidence for valid, appropriate score interpretations, including total-score and subscore reliability/precision, standard error of measurement, and inter-rater reliability for writing and speaking responses.","Analysis of the Scoring and Reliability for the Duolingo English Test  \nTable of Contents  \nSection  \n1 Purpose of this Document  \n2 Sources Reviewed  \n3 DET Scoring Procedures  \n4 Score Reliability/ Precision – Overview  \n5 Procedures for Item Generation and Estimation of Item-Level Difﬁculty  \n6 Procedures for Scoring Item Responses  \n7 Procedures for Deriving Total Scores and Subscores  \n8 Procedures for Supporting Valid and Appropriate Score Interpretations  \n9 DET Total Score Reliability/ Precision  \n10 DET Subscore Reliability/ Precision  \n11 Standard Error of Measurement  \n12 Inter-Rater Reliability of Writing and Speaking Responses  \nReferences  \nSection 1: Purpose of this Document  \nThe purpose of this document is to report our review of the research supporting the Duolingo English Test (DET) scoring and score reliability/precision. We used the Standards for Educational and Psychological Testing (AERA , APA , & NCME, 2014) and DET’s documentation to analyze the topics identiﬁed below. As the Standards address scoring and reliability issues across various chapters, this document integrates sections from chapters 2, 5, and 6. Key considerations related to scoring and reliability associated with the chapters include:  \nStandard 2.0: Appropriate evidence of reliability/precision should be provided for the interpretation for each intended score use (AERA , APA , & NCME, 2014, p. 42) .  \nStandard 5.0: Test scores should be derived in a way that supports the interpretations of test scores for the proposed uses of tests. Test developers and users should document evidence of fairness, reliability, and validity of test scores for their proposed use (AERA , APA , & NCME, 2014, p. 102) .  \nStandard 6.0: To support useful interpretations of score results, assessment instruments should have established procedures for test administration, scoring, reporting, and interpretation. Those responsible for administering, scoring, reporting, and interpreting should have sufﬁcient training and supports to help them follow the established procedures. Adherence to the established procedures should be monitored, and any material errors should be documented and, if possible, corrected (AERA , APA , & NCME, 2014, p. 114) .  \nQuestions Analyzed  \nWe analyzed the following questions related to the scoring of the DET:  \n1. What are the established scoring procedures for the DET?  \n● What are DET ’s procedures for generating items and estimating item-level difﬁculty?  \n● What are DET ’s procedures for selecting items for administration?  \n● What are DET ’s procedures for scoring item responses?  \n● What are DET ’s procedures for deriving total scores and subscores?  \n● What are DET ’s procedures for supporting valid and appropriate score interpretations?  \n2. What is the level of reliability/precision of the DET scores?  \n● What is the reliability/precision of the DET total scores?  \n● What is the reliability/precision of the DET subscores?  \n● What is the amount of error contained within the DET scores?  \nIn what follows, we summarize our analysis and use examples of evidence collected from DET ’s documentation to address the questions.  \nBack to Table of Contents  \nSection 2: Sources Reviewed  \nTo address the aforementioned goals and questions, we reviewed publicly-available DET documentation (e.g. , articles and DET websites for test takers and institutions) . We also reviewed a few independently completed evaluations of the DET and held semi-structured interviews with DET staff. The staff included the chief of assessment and assessment scientists. The published resources used in the development of this report are cited in the reference section at the end of this document.  \nBack to Table of Contents  \nSection 3: DET Scoring Procedures  \nStandard 5.16: When test scores are based on model-based psychometric procedures, such as those used in computerized adaptive or multistage testing, documentation should be provided to indicate that scores have co","cbCaidXvxfcPQVhZ","https://ap.wps.com/l/cbCaidXvxfcPQVhZ","pdf",346597,6,1,27,"English","en",105,"# Section 1 Purpose of this Document\n# Section 2 Sources Reviewed\n# Section 3 DET Scoring Procedures\n## Overview of Scoring Procedures\n## Computer-Adaptive Testing (CAT)","[{\"question\":\"What is the purpose of this document on the DET?\",\"answer\":\"It reports the review of research supporting DET scoring and the reliability/precision of scores. It evaluates scoring-related procedures and evidence for score interpretation.\"},{\"question\":\"How does the DET scoring process estimate item difficulty?\",\"answer\":\"DET uses machine learning and natural language processing models to create proficiency scales and to estimate item difficulty for computer-adaptive test (CAT) algorithms.\"},{\"question\":\"Which reliability and precision aspects does the document analyze?\",\"answer\":\"It covers DET total-score reliability/precision, subscore reliability/precision, the standard error of measurement, and inter-rater reliability for writing and speaking responses.\"}]","Analysis of the Scoring and Reliability for the Duolingo English Test - Procedures and Reliability/Precision | PDF",1788068530,68,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"analysis-of-the-scoring-and-reliability-for-the-duolingo-english-test-procedures-and-reliabilityprecision","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/exam/",3,{"item":53,"name":13,"@type":44,"position":11},"https://docshare.wps.com/document/analysis-of-the-scoring-and-reliability-for-the-duolingo-english-test-procedures-and-reliabilityprecision/160530/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-05","2026-08-30",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the purpose of this document on the DET?","Question",{"text":76,"@type":77},"It reports the review of research supporting DET scoring and the reliability/precision of scores. It evaluates scoring-related procedures and evidence for score interpretation.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the DET scoring process estimate item difficulty?",{"text":81,"@type":77},"DET uses machine learning and natural language processing models to create proficiency scales and to estimate item difficulty for computer-adaptive test (CAT) algorithms.",{"name":83,"@type":74,"acceptedAnswer":84},"Which reliability and precision aspects does the document analyze?",{"text":85,"@type":77},"It covers DET total-score reliability/precision, subscore reliability/precision, the standard error of measurement, and inter-rater reliability for writing and speaking responses.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,105,110,114,119,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":103,"slug":104},70,"exam",{"id":106,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":122,"slug":123},8,"Research & Report",30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":106,"slug":139},19,"General","general"]