[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82990-en":3,"doc-seo-82990-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82990,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","PERSONAJUDGE Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data","Large language models increasingly act as judges in AI evaluation, yet common methods rely on consensus preferences that overlook how individual evaluators differ. PERSONAJUDGE introduces a simulation approach that uses evaluator-specific auxiliary signals—retrospective reasoning traces and interface telemetry—combined with categorical judgments to enable LLM-based individual simulation via in-context learning. A systematic study uses 32 trained annotators and 4,200 judgments in a 4×4×4 factorial design, showing up to 9.9-point gains over a base judge, with reasoning traces helping most and telemetry often reducing performance.","PERSONAJUDGE: Simulating Individual Human Preference Judgments  \nwith Evaluator-Specific Demonstration Data  \nZeyu He1* , Xuan Qi2 , Subramanian Chidambaram2 , Zhichao Xu2 ,  \nVinayak Arannil2 , Lydia Chilton2,3 , Alex C. Williams2  \n1Pennsylvania State University, 2AWS AI Fundamental Research, 3 Columbia University  \nCorrespondence: [zeyuhe@psu.edu](zeyuhe@psu.edu)  \narXiv :2607 .05742v 1 [ cs .HC] 7 Jul 2026  \nAbstract  \nLarge language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data—retrospective reasoning traces and interface telemetry—to enable LLM-based simulation of individual evaluators via in-context learning. We conduct a systematic empirical study of this approach using multi-facet data from 32 trained annotators across 4,200 preference judgments in a 4 × 4 × 4 factorial design. Our key findings: (1) The simulation approach achieves up to 9.9 percentage point improvements over the Base Judge; (2) Reasoning traces provide the largest gains with higher collection efforts, while interface telemetry often hurts rather than helps performance despite being cheaper to collect. (3) Simulation difficulty is systematic, predicted by an evaluator’s neutral usage (most clearly on Helpfulness) and divergence from consensus; the neutral-usage tendency—rather than simulatability itself—is the cross-task-stable property (r = 0 .728) . These results establish both the potential and limits of evaluator-specific auxiliary data for personalized evaluation, offering methodological insights for scaling individualaware AI assessment.  \n1 Introduction  \nLarge Language Models (LLMs) are increasingly used as judges for model comparison, reward modeling, and benchmark construction because they offer a scalable alternative to human evaluation (Bai et al., 2022b ; Liu et al., 2023 ; Zheng et al., 2023 ; Lambert et al., 2025) . However, most LLM-asJudge pipelines target an aggregate signal: models are trained, aligned, or evaluated against pooled  \n*  \nWork completed during ZH’s Amazon internship.  \nFigure 1: PERSONAJUDGE workflow. A human evaluator first completes judgment tasks, producing three complementary signals: a categorical judgment, interface telemetry, and retrospective reasoning. PERSONAJUDGE organizes these signals into evaluator-specific demonstrations and combines them with the task instruction and a new task instance. An LLM then uses this information to simulate the target evaluator’s judgment.  \npreferences that collapse across annotators (Stiennon et al., 2020 ; Ouyang et al., 2022 ; Bai et al., 2022a ; Zhang et al., 2025), so they simulate crowdlevel consensus rather than any specific evaluator. This is a key limitation for open-ended evaluation, where a single shared standard may not exist and annotators apply different criteria to the same outputs (Van der Lee et al., 2021 ; Davani et al., 2022) . In such settings disagreement is often meaningful rather than noise (Basile et al., 2021 ; Fleisiget al., 2023 ; Frenda et al., 2025) . Even LLM judges that match average ratings can miss group- and person-level differences (Movva et al., 2024) . Current models struggle to infer individual preferences from value-laden statements (Jiang et al., 2025) . Matching consensus thus does not guarantee faithful simulation of any individual evaluator.  \nThis motivates a different target: individual evaluator simulation, which can support evaluation auditing, per-person reward modeling, and fairness analysis of whose perspectives evaluation pipelines underrepresent. Prior personalized methods condition on persona descriptions or compact preference  \nprofiles (Dong et al., 2024b ; Wang et al., 2024a), but represent people through summaries rather than their own decision traces. We instead ask: given an evaluator’s historical judgm","cbCaieSH7g0oxzfu","https://ap.wps.com/l/cbCaieSH7g0oxzfu","pdf",3820736,4,1,22,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does PERSONAJUDGE address in LLM-as-judge evaluation?\",\"answer\":\"Most pipelines target aggregate consensus, which can fail to reflect individual evaluator criteria in open-ended settings where annotators disagree meaningfully. PERSONAJUDGE targets individual evaluator simulation instead of crowd-level consensus.\"},{\"question\":\"How does PERSONAJUDGE simulate an individual evaluator’s judgments?\",\"answer\":\"It constructs evaluator-specific demonstrations from multiple signals: a categorical judgment, interface telemetry, and retrospective reasoning. These are combined with task instructions and a new instance so an LLM can simulate the target evaluator via in-context learning.\"},{\"question\":\"Which auxiliary data source improves simulation the most, and which can hurt performance?\",\"answer\":\"Retrospective reasoning traces provide the largest gains, especially with higher collection effort. Interface telemetry often hurts performance despite being cheaper to collect, according to the empirical findings.\"}]",1784184497,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"personajudge-simulating-individual-human-preference-judgments-with-evaluator-specific-demonstration-data","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/personajudge-simulating-individual-human-preference-judgments-with-evaluator-specific-demonstration-data/82990/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does PERSONAJUDGE address in LLM-as-judge evaluation?","Question",{"text":75,"@type":76},"Most pipelines target aggregate consensus, which can fail to reflect individual evaluator criteria in open-ended settings where annotators disagree meaningfully. PERSONAJUDGE targets individual evaluator simulation instead of crowd-level consensus.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does PERSONAJUDGE simulate an individual evaluator’s judgments?",{"text":80,"@type":76},"It constructs evaluator-specific demonstrations from multiple signals: a categorical judgment, interface telemetry, and retrospective reasoning. These are combined with task instructions and a new instance so an LLM can simulate the target evaluator via in-context learning.",{"name":82,"@type":73,"acceptedAnswer":83},"Which auxiliary data source improves simulation the most, and which can hurt performance?",{"text":84,"@type":76},"Retrospective reasoning traces provide the largest gains, especially with higher collection effort. Interface telemetry often hurts performance despite being cheaper to collect, according to the empirical findings.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]