[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83647-en":3,"doc-seo-83647-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83647,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Personality Without Persons A Psychometric Critique of Big Five Testing in Large Language Models","Human personality inventories increasingly characterize large language models (LLMs), support system comparisons, and underpin downstream governance claims. However, these instruments were built and validated for humans, leaving uncertainty about their applicability to LLMs. A systematic psychometric evaluation tests Big Five personality measurements in LLMs by examining content validity across multiple inventories and administering the top inventory to N=244 models from 49 families. Results show limited inter-model variability, failure to recover the five-factor structure, and shifts toward socially desirable traits after alignment training, undermining the equivalence to human personality.","Personality Without Persons?  \nA Psychometric Critique of Big Five Testing in Large Language Models  \nKim Zierahn 1 , Cristina Cachero2 , Anna Korhonen3 , Nuria Oliver 1  \n1ELLIS Alicante, Spain  \n2University of Alicante, Spain  \n3University of Cambridge, United Kingdom  \n[kim@ellisalicante.org](kim@ellisalicante.org), [ccachero@dlsi.ua.es](ccachero@dlsi.ua.es), [alk23@cam.ac.uk](alk23@cam.ac.uk), [nuria@ellisalicante.org](nuria@ellisalicante.org)  \narXiv :2607 .02325v 1 [ cs .HC] 2 Jul 2026  \nAbstract  \nHuman personality inventories are increasingly used to characterize large language models (LLMs), compare systems, and inform downstream governance claims. Yet, these inventories were developed and validated for humans, and it remains unclear whether they apply to LLMs. We present a systematic psychometric evaluation of Big Five personality measurements in LLMs. We ask three research questions: Do Big Five inventories a) appropriately describe LLMs, b) capture inter-individual differences across models, and c) reﬂect internal factors consistent with human personality? We assess content validity of ﬁve candidate Big Five inventories and administer the winning inventory to N = 244 different models spanning 49 model families. First, we found that Big Five items adapted for LLMs can reach sufﬁcient content validity, while original human-developed items do not. Second, Big Five inventories do not capture meaningful differences between LLMs: We found low variability between models, accounting for only 3% of total score variance. Third, LLMs responses do not recover the Big Five ﬁve-factor structure with four of the Big Five facets collapsing into one (r ≥ .92) . Direct comparisons between base and instruction-tuned model variants suggest that alignment training systematically shifts Big Five scores toward socially desirable traits. These ﬁndings demonstrate that Big Five scores do not measure a construct equivalent to human personality in LLMs. Applying human personality frameworks to LLMs produces misleading characterizations used to benchmark, compare, and govern LLMs. We highlight the need for evaluation frameworks that are developed for LLMs, rather than adopting human constructs without validation.  \nData and code—[https://github.com/ellisalicante/BigFive](https://github.com/ellisalicante/BigFive)LLM-Evaluation  \n1 Introduction  \nLarge language models (LLMs) no longer serve just as tools. Users often attribute the role of an advisor, assistant or companion to ChatGPT, covering topics such as health, ﬁnance, role-play, mental health, and programming (Karnam et al. 2026) . These roles require LLMs to be servile and helpful, act as role models, and avoid toxic outputs (Bodroˇza, Dinic´, and Boji´c 2024; Lee et al. 2025; Li et al. 2024) . LLMs can  \nCopyright © 2026, Association for the Advancement of Artiﬁcial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \ncreate the illusion of human-likeness and even “personality”, by producing the presence of social cues (Nass and Moon 2000; Fogg 2003; Peter, Riemer, and West 2025) . At the same time, both users and companies describe chatbots using human categories (Nass and Moon 2000; Peter, Riemer, and West 2025): some models appear cautious and polite; others conﬁdent and direct; some empathetic and encouraging. These perceived traits can inﬂuence users’ trust, persuasion, engagement, motivation, and enjoyment (Kovaˇcevi´cet al. 2024; Kuhail et al. 2024; Lee et al. 2024; Moilanenet al. 2022; S¨oderqvist 2025; Sonlu et al. 2024) .  \nYet the scientiﬁc concept of whether LLMs truly exhibit personality-like characteristics remains deeply contested. Treating LLMs as if they have personality risks misleading assumptions about consciousness, feeling, and intention, and risks persuasion and manipulation (Peter, Riemer, and West 2025) . Nass and Moon (2000) already found 25 years ago that users mindlessly apply social expectations and human stereotypes to computers. With the cur","cbCaiefLg6rpkjWX","https://ap.wps.com/l/cbCaiefLg6rpkjWX","pdf",4567178,3,1,19,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Methodology\n# Findings\n# Implications for Governance\n# Data and Code","[{\"question\":\"What research questions does the paper investigate about Big Five testing in LLMs?\",\"answer\":\"It asks whether Big Five inventories appropriately describe LLMs, whether they capture meaningful inter-model differences, and whether LLM responses reflect internal factors consistent with the human Big Five structure.\"},{\"question\":\"How is the Big Five inventory evaluated across models in the study?\",\"answer\":\"The study evaluates content validity for five candidate Big Five inventories, selects the winning inventory based on expert content validity, and administers it to N=244 LLMs spanning 49 model families.\"},{\"question\":\"What are the main results regarding whether Big Five scores measure human-equivalent personality in LLMs?\",\"answer\":\"LLM-adapted items reach sufficient content validity, but Big Five inventories show low variability across models and do not recover the expected five-factor structure; additionally, instruction tuning can shift scores toward socially desirable traits.\"}]",1784189499,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"personality-without-persons-a-psychometric-critique-of-big-five-testing-in-large-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/personality-without-persons-a-psychometric-critique-of-big-five-testing-in-large-language-models/83647/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What research questions does the paper investigate about Big Five testing in LLMs?","Question",{"text":75,"@type":76},"It asks whether Big Five inventories appropriately describe LLMs, whether they capture meaningful inter-model differences, and whether LLM responses reflect internal factors consistent with the human Big Five structure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the Big Five inventory evaluated across models in the study?",{"text":80,"@type":76},"The study evaluates content validity for five candidate Big Five inventories, selects the winning inventory based on expert content validity, and administers it to N=244 LLMs spanning 49 model families.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main results regarding whether Big Five scores measure human-equivalent personality in LLMs?",{"text":84,"@type":76},"LLM-adapted items reach sufficient content validity, but Big Five inventories show low variability across models and do not recover the expected five-factor structure; additionally, instruction tuning can shift scores toward socially desirable traits.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]