[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83699-en":3,"doc-seo-83699-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83699,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Silicon Sampling via Cross-Survey Transfer","Silicon sampling uses large language models (LLMs) to simulate human survey respondents, but many evaluations focus on distributional similarity rather than individual-level prediction. A cross-survey transfer framework is introduced: the model is conditioned on one set of answers and must predict responses to entirely different items. Using TEDS 2024 data, three open-weight LLMs and supervised baselines show strong zero-shot accuracy, a stable predictability hierarchy, and nuanced effects of variance collapse and safety alignment.","Silicon Sampling via Cross-Survey Transfer  \nChan-Tung Ku  \nDepartment of Information Management National Sun Yat-sen University Kaohsiung, Taiwan [kuchantung@gmail.com](kuchantung@gmail.com)  \nChan Hsu  \nDepartment of Information Management National Sun Yat-sen University Kaohsiung, Taiwan [chanshsu@gmail.com](chanshsu@gmail.com)  \nPei-Cing Huang  \nDepartment of Information Management National Sun Yat-sen University Kaohsiung, Taiwan [pcpeicing@gmail.com](pcpeicing@gmail.com)  \nFrank Cheng-shan Liu Institute of Political Science National Sun Yat-sen University Kaohsiung, Taiwan [csliu@mail.nsysu.edu.tw](csliu@mail.nsysu.edu.tw)  \nI-Ling Cheng  \nGraduate Institute of Library and Information Science National Chung Hsing University Taichung, Taiwan [chengi428@gmail.com](chengi428@gmail.com)  \nYihuang Kang  \nDepartment of Information Management National Sun Yat-sen University Kaohsiung, Taiwan [ykang@mis.nsysu.edu.tw](ykang@mis.nsysu.edu.tw)  \nAbstract—Silicon sampling—using large language models (LLMs) to simulate human survey respondents—has emerged asa promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent’s answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B–120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects—two commonly cited LLM limitations—turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.  \nKeywords—silicon sampling, large language models, survey simulation, cross-survey transfer, political attitudes, social simulation  \nI. INTRODUCTION  \nSurvey research faces growing practical challenges. Response rates have declined substantially over recent decades, while costs per completed interview continue to rise [1] . At the same time, policymakers and researchers increasingly need rapid, fine-grained measurements of public opinion—precisely when traditional methods are becoming slower and more expensive.  \nIn this context, the idea of using large language models (LLMs) as simulated survey respondents—termed silicon sampling by Argyle et al. [2]—has attracted considerable attention. The premise is straightforward: condition an LLM on a demographic profile and attitudinal context, then ask it to complete a survey as that persona. If the simulated responses  \napproximate real ones, silicon sampling could dramatically reduce the cost and turnaround time of survey research. Early results have been encouraging, with LLMs reproducing aggregate opinion distributions from real surveys [2] and predicting experimental treatment effects with striking accuracy [3] .  \nHowever, a methodological concern undermines much of this optimism. Most evaluations condition LLMs on demographic profiles and assess performance at the distributional level—comparing aggregate response distributions between silicon and human samples [2], [4], [5] . Even the most rigorous designs, such as Argyle et al.’s Study 3 (which uses 11 survey items to predict a 12th), evaluate distributional correspondence rather than individual-level prediction accu","cbCaiuH3eGQvNSIU","https://ap.wps.com/l/cbCaiuH3eGQvNSIU","pdf",464888,4,1,6,"English","en",105,"# Introduction\n## Motivation and challenges in survey research\n## Silicon sampling and distributional evaluation limits\n## Cross-survey transfer as an individual-level test\n# Method and evaluation setup\n## Dataset: TEDS 2024 and target constructs\n## Models and supervised baselines\n## Alignment-removal and study research questions","[{\"question\":\"What is cross-survey transfer in silicon sampling?\",\"answer\":\"Cross-survey transfer partitions survey items into context (Set A) and prediction (Set B) sets. An LLM receives a respondent’s answers to Set A and must predict their answers to different items in Set B.\"},{\"question\":\"How does this study evaluate LLM performance compared with prior work?\",\"answer\":\"Rather than relying on aggregate distribution matching, the study tests individual-level generalization by predicting genuinely unseen items. It compares zero-shot LLMs to supervised machine learning baselines trained on same-population data.\"},{\"question\":\"What did the findings reveal about construct predictability and known LLM limitations?\",\"answer\":\"The results show a stable construct predictability hierarchy across different survey topics. They also find variance collapse can affect supervised models and that safety alignment effects vary substantially across model families, refining earlier claims about these limitations.\"}]",1784189808,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"silicon-sampling-via-cross-survey-transfer","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/silicon-sampling-via-cross-survey-transfer/83699/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is cross-survey transfer in silicon sampling?","Question",{"text":75,"@type":76},"Cross-survey transfer partitions survey items into context (Set A) and prediction (Set B) sets. An LLM receives a respondent’s answers to Set A and must predict their answers to different items in Set B.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does this study evaluate LLM performance compared with prior work?",{"text":80,"@type":76},"Rather than relying on aggregate distribution matching, the study tests individual-level generalization by predicting genuinely unseen items. It compares zero-shot LLMs to supervised machine learning baselines trained on same-population data.",{"name":82,"@type":73,"acceptedAnswer":83},"What did the findings reveal about construct predictability and known LLM limitations?",{"text":84,"@type":76},"The results show a stable construct predictability hierarchy across different survey topics. They also find variance collapse can affect supervised models and that safety alignment effects vary substantially across model families, refining earlier claims about these limitations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]