[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83302-en":3,"doc-seo-83302-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83302,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Validating LLMs in Social Science: Epistemic Threats and Emerging Norms","Large language models (LLMs) are increasingly used in social science to generate quantitative measurements of social concepts, such as annotating data, simulating survey responses, and estimating ideology. Yet bias, hallucination, and context brittleness create unclear threats to validity, while field-wide standards for validation remain underdeveloped. This study compiles and systematically analyzes validation practices across a corpus from eight major social science journals, finding central use of LLM measurements paired with limited and inconsistent validation, and proposes complementary strategies for more robust norms.","arXiv :2607 .079 15v 1 [ cs .CY] 8 Jul 2026  \nValidating LLMs in social science: Epistemic threats  \nand emerging norms  \nMeera Desai 1∗ , Dallas Card 1 , Abigail Z. Jacobs 1,2  \n1 School of Information, University of Michigan, Ann Arbor, USA  \n2 Center for the Study of Complex Systems, University of Michigan, Ann Arbor, USA.  \n∗ [Corresponding author. Email: madesai@umich.edu](Corresponding author. Email: madesai@umich.edu)  \nLarge language models (LLMs) are reshaping social science methodology.  \nResearchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling data or simulating survey responses. Yet LLMs pose methodological challenges including bias, hallucination, and brittleness across contexts, with unclear threats to validity. Standard practices and norms for addressing these challenges are still emerging. We collect and systematically analyze validation practices in a comprehensive corpus of papers from eight flagship social science journals that use LLMs as measurement instruments. We find that LLM-generated measurements frequently play a central role in empirical analyses, yet validation practices are inconsistent and limited.  \nWe outline complementary strategies for more robust validation, pointing toward better norms and standards around the use of LLMs in social science.  \n1 Introduction  \nAcross fields, social scientists are increasingly adopting the practice of prompting large language models (LLMs) to generate quantitative measures of social concepts. For example, researchers use LLMs as measurement instruments by prompting LLMs to annotate and code data, simulate survey responses, and estimate ideological positions. However, the progression and legitimacy of social science rely on the ability to make valid arguments. In turn, making valid arguments often rests on the ability to take valid measurements of social constructs, like ideology, emotion, or sentiment. This paper empirically examines the recent trend of using LLMs as measurement instruments and identifies important validity challenges (and opportunities) for social science research.  \nThe widespread adoption of LLMs as measurement instruments reflects their broad appeal as a faster, cheaper, and potentially more accurate alternative to humans at tedious tasks like annotating or coding data (1,2) . Researchers also argue that LLMs are easier to use than alternative computational approaches (such as training bespoke machine learning models; (3)), and that their performance and flexibility makes them applicable to a wide range of complex tasks than earlier computational text analysis methods embraced by social scientists (4, 5) .  \nWhether researchers use LLMs, dictionary-based approaches, human coding, or any other measurement instrument, measuring abstract constructs requires making interpretive choices about a concept’s definition and how it can be observed, leaving room for disagreement and error (6, 7) . These methodological choices are consequential and contestable, and validity cannot be assumed. Researchers have always had to develop and engage with contextually appropriate ways of validating their measurements even as computational tools have become increasingly sophisticated (8, 9, 10) . Despite the potential of LLMs as efficient tools for measuring social concepts, there are plenty of reasons to be skeptical about using LLMs as measurement instruments. Research from natural language processing has highlighted many challenges associated with LLMs including their sensitivity to small changes in prompt or configuration (11, 12, 13, 14, 15, 16, 17), bias (18, 19), and poor calibration (20, 21, 22) . Along with other model limitations, these issues can affect the factual reliability of LLMs, which sometimes produce “hallucinated” outputs that are coherent but incorrect. Moreover, errors are difficult to predict or mitigate, as neither the precise nature of model biases nor the relationsh","cbCairBu8aRdJKhh","https://ap.wps.com/l/cbCairBu8aRdJKhh","pdf",681651,1,28,"English","en",105,"# Introduction\n# Results\n## Overview and use cases","[{\"question\":\"LLMs如何在社会科学研究中被用作“测量工具”？\",\"answer\":\"研究者通过提示LLM进行数据标注与编码、模拟调查回应，以及估计意识形态等方式，把抽象社会概念转化为量化测量。\"},{\"question\":\"在使用LLMs进行测量时，主要的有效性威胁有哪些？\",\"answer\":\"文中强调了偏差、幻觉输出以及在不同情境中的脆弱性，这些因素使得威胁到效度的具体问题仍不清晰。\"},{\"question\":\"作者对当前验证实践的主要结论是什么，并提出了什么改进方向？\",\"answer\":\"作者发现LLM生成的测量常常在实证分析中处于核心地位，但验证实践有限且不一致；因此提出互补的更稳健验证策略，以推动更好的规范与标准。\"}]",1784186624,71,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"validating-llms-in-social-science-epistemic-threats-and-emerging-norms","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/validating-llms-in-social-science-epistemic-threats-and-emerging-norms/83302/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"LLMs如何在社会科学研究中被用作“测量工具”？","Question",{"text":75,"@type":76},"研究者通过提示LLM进行数据标注与编码、模拟调查回应，以及估计意识形态等方式，把抽象社会概念转化为量化测量。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"在使用LLMs进行测量时，主要的有效性威胁有哪些？",{"text":80,"@type":76},"文中强调了偏差、幻觉输出以及在不同情境中的脆弱性，这些因素使得威胁到效度的具体问题仍不清晰。",{"name":82,"@type":73,"acceptedAnswer":83},"作者对当前验证实践的主要结论是什么，并提出了什么改进方向？",{"text":84,"@type":76},"作者发现LLM生成的测量常常在实证分析中处于核心地位，但验证实践有限且不一致；因此提出互补的更稳健验证策略，以推动更好的规范与标准。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]