[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126765-en":3,"doc-seo-126765-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126765,962084926284,"Aurora","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Moving to continuous classifications of bilingualism through machine learning trained on language production","Recent conceptualisations of bilingualism move away from strict categorisations toward continuous approaches. This study combines empirical psycholinguistics data with machine learning classification modelling. Support vector classifiers were trained on two datasets of linguistic productions by Italian speakers to predict classes (monolingual, attriters, heritage). All classes were identified above chance, with monolinguals best and attriters most confusable. Confusion patterns indicate attriters often align with heritage speakers, suggesting a middle position. Clitic clusters proved especially informative, supporting bilingualism as a continuum of linguistic behaviours.","1 Short title: Moving to continuous classifications of bilingualism through machine  \n2 learning 3  \n4 Long title: Moving to continuous classifications of bilingualism through machine  \n5 learning trained on language production 6  \n7 Authors: Coco, M. I. 1,2*, Smith, G.3*, Spelorzi, R.4 & Garraffa, M.3  \n8 1Department of Psychology,“Sapienza” University of Rome, Rome, Italy  \n9 2 I.R.C. S. S Fondazione Santa Lucia, Rome, Italy  \n10 3 School of Health Sciences, University of East Anglia, Norwich, UK  \n11 4Department of Linguistics and English Language, University of Edinburgh, Edinburgh, UK 12  \n13 * Denotes equal contribution  \n14  \n15 Competing interests: the authors declare none.  \n16  \n17 Addresses for correspondence: 18  \n19 Dr Moreno I. Coco  \n20 Dipartimento di Psicologia  \n21 Sapienza, Universita’ di Roma  \n22 Via dei Marsi, 78  \n23 00185, Roma,  \n24 Italy  \n25 [email: ](email: moreno.coco@uniroma1.it)[moreno.coco@uniroma1.it](email: moreno.coco@uniroma1.it)[ ](email: moreno.coco@uniroma1.it)26  \n27 Dr Giuditta Smith  \n28 School of Health Sciences  \n29 University of East Anglia  \n30 Norwich Research Park  \n31 NR4 7TJ  \n32 Norwich  \n33 UK  \n34 Email: [giuditta.smith@uea.ac.uk](giuditta.smith@uea.ac.uk)  \n35  \n36  \n37 Abstract  \n38  \n39 Recent conceptualisations of bilingualism are moving away from strict categorisations, 40 towards continuous approaches. This study supports this trend by combining empirical  \n41 psycholinguistics data with machine learning classification modelling. We trained support  \n42 vector classifiers on two datasets of linguistic productions coded for type of production of  \n43 Italian speakers to predict their class (i.e., “monolingual”, “attriters”, and “heritage”) . All  \n44 classes were predicted above chance (> 33%), even if the classifier’s performance substantially  \n45 varies, with monolinguals identified better (f-score > 70%) and attriters being the most  \n46 confusable (f-score \u003C 50%) . The confusion matrices qualify that attriters are identified as  \n47 heritage speakers nearly as often as they could be correctly classified, suggesting this class to  \n48 sit in the middle. Clitic clusters were found to be the most identifying features for  \n49 discrimination. Overall, this study supports a conceptualisation of bilingualism as a continuum  \n50 of linguistic behaviours rather than sets of a-priori established classes.  \n51 Keywords: bilingualism, heritage speakers, attrition, support vector machine, classification  \n52  \n53 1. Introduction  \n54  \n55 In a globalized and highly integrated world, the boundaries of languages have become fluid  \n56 and seemingly continuous. Speakers are more likely to move across countries, transfer their  \n57 homeland language to their offspring and acquire other languages, with bilingual proficiency  \n58 reaching native-like language abilities well after childhood (Steinhauer, 2014; Roncaglia- 59 Denissen & Kotz, 2016; Hartshorne, Tenenbaum, & Pinker, 2018; Köpke, 2021; Gallo et al.  \n60 2021). However, bilingualism is known to substantially vary among individuals, as it is  \n61 shaped by intra and extralinguistic factors such as amount of exposure, social status, and  \n62 education (Gullifer et al. 2018; Haranto & Yang, 2016; Rodina et al. 2020; Bialystok, 2016;  \n63 Gullifer & Titone, 2020) . Consequently, research in bilingualism has progressively  \n64 abandoned strict categorical approaches in favour of more nuanced ones. The increased  \n65 complexity of a “winner-take-all” definition of bilingualism has created a plethora of labels  \n66 to classify speakers (see Surrain & Luk, 2017 for a systematic review), sometimes leading to  \n67 the same speakers being labelled differently according to whether the classification is based  \n68 on language dominance, learning history, age, etc., which renders it impractical to perform  \n69 consistent comparisons across different studies. Moreover, strict classifications disregard that  \n70 individuals can also change","cbCaitkihmO7XfFB","https://ap.wps.com/l/cbCaitkihmO7XfFB","pdf",563076,1,34,"English","en",105,"# Abstract\n# Introduction\n## Background and limitations of categorical bilingualism\n## Research question and implications for terminology and methodology","[{\"question\":\"What problem does the study address in bilingualism research?\",\"answer\":\"It addresses how strict categorical labels for bilingualism may overlap and fail to support consistent comparisons across studies, since individuals can change language-related profiles over time.\"},{\"question\":\"How were machine learning models used in this research?\",\"answer\":\"Support vector classifiers were trained on two datasets of Italian speakers’ linguistic productions to predict classes: monolingual, attriters, and heritage.\"},{\"question\":\"Which linguistic features and class relationships were most informative?\",\"answer\":\"Clitic clusters were identified as the most discriminative features. The confusion matrix showed attriters were often classified as heritage speakers, indicating they may lie between the two categories.\"}]","Moving to continuous classifications of bilingualism through machine learning trained on language production | PDF",1785934668,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"moving-to-continuous-classifications-of-bilingualism-through-machine-learning-trained-on-language-production","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/moving-to-continuous-classifications-of-bilingualism-through-machine-learning-trained-on-language-production/126765/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in bilingualism research?","Question",{"text":75,"@type":76},"It addresses how strict categorical labels for bilingualism may overlap and fail to support consistent comparisons across studies, since individuals can change language-related profiles over time.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How were machine learning models used in this research?",{"text":80,"@type":76},"Support vector classifiers were trained on two datasets of Italian speakers’ linguistic productions to predict classes: monolingual, attriters, and heritage.",{"name":82,"@type":73,"acceptedAnswer":83},"Which linguistic features and class relationships were most informative?",{"text":84,"@type":76},"Clitic clusters were identified as the most discriminative features. The confusion matrix showed attriters were often classified as heritage speakers, indicating they may lie between the two categories.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]