[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-473187-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-473187-en":130},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","corpus-based-identification-and-profiling-of-vocabulary-in-linguaskill-speaking-tests-a-methodological-framework-for-adaptive-test-validation","Corpus-Based Identification and Profiling of Vocabulary in Linguaskill Speaking Tests - A Methodological Framework for Adaptive Test Validation","","Computer-adaptive tests such as Cambridge Linguaskill adjust item difficulty dynamically to match test-taker proficiency, and their speaking components are now used for high-stakes admission and exit decisions, including in Malaysian higher education. Yet little research has examined the vocabulary load and lexical demands of adaptive speaking assessment, leaving open whether the vocabulary elicited at each reported level aligns with the Common European Framework of Reference (CEFR). This article develops a corpus-based framework for vocabulary identification, classification, and CEFR-aligned profiling. It specifies corpus compilation across CEFR strata, multi-lens analysis via LexTutor, New General Service List coverage, and Text Inspector against the English Vocabulary Profile, and decision rules to judge alignment, over- or under-demanding levels using spoken-English threshold benchmarks. The protocol is grounded in argument-based validation and addresses washback, fairness, and pedagogical stakes.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/corpus-based-identification-and-profiling-of-vocabulary-in-linguaskill-speaking-tests-a-methodological-framework-for-adaptive-test-validation/473187/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/corpus-based-identification-and-profiling-of-vocabulary-in-linguaskill-speaking-tests-a-methodological-framework-for-adaptive-test-validation/473187.png","ImageObject",300,407,{"name":42,"@type":43},"Olivia Brown","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-07","2026-09-30",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",5,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What issue does the article address in Linguaskill speaking assessments?","Question",{"text":63,"@type":64},"It examines whether the vocabulary elicited by adaptive speaking tasks is congruent with the CEFR levels claimed in score reports, despite limited prior scrutiny of lexical demands in adaptive speaking.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does the proposed framework identify and profile vocabulary?",{"text":68,"@type":64},"It compiles a specialized corpus of transcribed candidate responses stratified across CEFR levels, then applies a three-lens procedure using LexTutor for word-family analysis, the New General Service List for high-frequency lemma coverage, and Text Inspector for CEFR-aligned profiling against the English Vocabulary Profile.",{"name":70,"@type":61,"acceptedAnswer":71},"What decision criteria are defined for CEFR alignment?",{"text":72,"@type":64},"Operational criteria distinguish core versus peripheral vocabulary at each adaptive level, and decision rules classify each level as aligned, overdemanding, or under-demanding relative to CEFR expectations using benchmark lexical coverage thresholds for spoken English.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},473187,1790815503,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,100,105,110,114,119,122,126],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":97,"show_sort_weight":98,"slug":99},"Comic",60,"comic",{"id":101,"doc_module":4,"doc_module_name":25,"category_name":102,"show_sort_weight":103,"slug":104},6,"Technology",50,"technology",{"id":106,"doc_module":4,"doc_module_name":25,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":111,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":112,"slug":113},8,30,"research-report",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":25,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":25,"category_name":128,"show_sort_weight":55,"slug":129},19,"General","general",{"code":4,"msg":82,"data":131},{"doc_id":79,"user_id":132,"nickname":42,"user_avatar":133,"doc_module":4,"category_id":111,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":139,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":144,"read_time":145},16904993612988,"https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd","Corpus-Based Identification and Profiling of Vocabulary in Linguaskill Speaking Tests: A Methodological Framework for  \nAdaptive Test Validation  \nRosatiqah Yahya1, Ng Yu Jin1*, Mallika Vasugi A/P V. Govindarajoo2, Kew Si Na3 and Pipit Rahayu4  \n1UNITAR University College Kuala Lumpur, Malaysia  \n2UNITAR International University, Malaysia  \n3Universitiy Teknologi Malaysia, Malaysia  \n4Universitas Pasir Pengaraian, Indonesia  \n*Corresponding Author  \nDOI: [https://doi.org/10.47772/IJRISS.2026.100701181](https://doi.org/10.47772/IJRISS.2026.100701181)  \nReceived: 26 July 2026; Accepted: 31 July 2026; Published: 24 August 2026  \nABSTRACT  \nComputer-adaptive tests such as Cambridge Linguaskill adjust item difficulty dynamically to match test-taker proficiency, and their speaking components are now used for high-stakes admission and exit decisions, including in Malaysian higher education. Yet little research has examined the vocabulary load and lexical demands of adaptive speaking assessment, leaving open whether the vocabulary elicited at each reported level is congruent with the Common European Framework of Reference (CEFR) levels the scores claim to represent. This article develops a complete corpus-based methodological framework for identifying, classifying, and profiling the vocabulary of the Linguaskill Speaking Test. The framework specifies the compilation of a specialised corpus of transcribed candidate responses stratified across CEFR levels, supplemented by official test preparation materials, and a three-lens profiling procedure using LexTutor for BNC-COCA word-family analysis, the New General Service List for high-frequency lemma coverage, and Text Inspector for CEFR-aligned profiling against the English Vocabulary Profile. Operational criteria are defined for distinguishing core from peripheral vocabulary at each adaptive level, and decision rules are specified for classifying each level as aligned, overdemanding, or under-demanding relative to CEFR expectations, benchmarked against published lexical coverage thresholds for spoken English. The central hypothesis, motivated by prior coverage research, is that the lexical demands elicited by adaptive speaking tasks will not rise in step with claimed CEFR levels. The framework is grounded in an argument-based approach to validation and articulates the washback, fairness, and pedagogical stakes of the alignment question. The article contributes a replicable protocol for lexical validation of adaptive speaking tests and a benchmark synthesis for interpreting its future results.  \nKeywords: Linguaskill speaking vocabulary; LexTutor; Text Inspector; corpus-based approach; vocabulary threshold; CEFR alignment; adaptive assessment  \nINTRODUCTION  \nLanguage testing has moved decisively online, and with it has come the rise of computer-adaptive assessment. Tests such as Cambridge Linguaskill adjust the difficulty of what candidates encounter in real time, tailoring the measurement to the individual rather than administering a fixed form to all (Chapelle & Voss, 2016) . The efficiency gains are considerable: shorter tests, faster results, and score reports mapped directly onto the levels of the Common European Framework of Reference for Languages (CEFR) . These properties have carried adaptive tests into high-stakes use. In Malaysia, Linguaskill has been recognised by the Ministry of Higher Education for university admission purposes since 2020 and is used by institutions for admission benchmarking  \nPage 17207  \n[www.rsisinternational.org](www.rsisinternational.org)  \nand exit requirements alongside the Malaysian University English Test, a context in which the meaning of a reported CEFR level carries direct consequences for students' progression.  \nA reported CEFR level is a claim, and claims require validation. The argument-based approach to test validation holds that every link in the chain from performance to score to interpretation to use must be supported by evidence (Kane, ","cbCaiqnoCQnpF8BF","https://ap.wps.com/l/cbCaiqnoCQnpF8BF","pdf",384546,18,"English","# Introduction\n## Adaptive language testing and CEFR-linked scoring\n## Argument-based validation and the vocabulary link\n## Lexical demand profiling under adaptivity","[{\"question\":\"What issue does the article address in Linguaskill speaking assessments?\",\"answer\":\"It examines whether the vocabulary elicited by adaptive speaking tasks is congruent with the CEFR levels claimed in score reports, despite limited prior scrutiny of lexical demands in adaptive speaking.\"},{\"question\":\"How does the proposed framework identify and profile vocabulary?\",\"answer\":\"It compiles a specialized corpus of transcribed candidate responses stratified across CEFR levels, then applies a three-lens procedure using LexTutor for word-family analysis, the New General Service List for high-frequency lemma coverage, and Text Inspector for CEFR-aligned profiling against the English Vocabulary Profile.\"},{\"question\":\"What decision criteria are defined for CEFR alignment?\",\"answer\":\"Operational criteria distinguish core versus peripheral vocabulary at each adaptive level, and decision rules classify each level as aligned, overdemanding, or under-demanding relative to CEFR expectations using benchmark lexical coverage thresholds for spoken English.\"}]","Corpus-Based Identification and Profiling of Vocabulary in Linguaskill Speaking Tests - A Methodological Framework for Adaptive Test Validation | PDF",1790782047,45]