[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81826-en":3,"doc-seo-81826-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81826,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Scaling Trends for Lie Detector Oversight in Preference Learning","Deceptive behavior in large language models is expensive to monitor and prevent, motivating scalable oversight methods such as Scalable Oversight via Lie Detectors (SOLiD), which routes likely deceptive responses for review by high-cost labelers. This work scales SOLiD to much larger models and tests it in more realistic preference-learning setups. Undetected deception improves with scale, dropping from 34% to 14% for 1B to 405B parameters at 99% detector true positive rate. Human labelers can be removed during finetuning without a significant deception increase. However, performance is sensitive to detector–data distribution shift, which can raise false positives to impractical levels.","Scaling Trends for Lie Detector Oversight in Preference Learning  \nOskar J. Hollinsworth* 1 Ann-Kathrin Dombrowski* 1 Sam Adam-Day 1 Adam Gleave 1 Chris Cundy 1  \narXiv :2607 .0 1567v 1 [ cs .AI] 2 Jul 2026  \nAbstract  \nDeceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detectors to identify responses for review by high-cost labelers. In this paper, we scale SOLiD to larger models and evaluate it in more diverse and realistic preference-learning settings.  \nWe find favorable scaling: undetected deception drops from 34% for 1B-parameter models to 14% for 405B-parameter models at a detector true positive rate of 99%, and expensive human labelers can be removed entirely from the finetuning phase without a statistically significant increase in deception. However, SOLiD is sensitive to distribution shift between detector training and preference-training data, which can drive detector false positive rates to impractical levels.  \n1. Introduction  \nEnsuring that AI systems pursue the behavior we want, rather than merely appearing to do so, is a central challenge for the field. Post-training techniques like Reinforcement Learning from Human Feedback (RLHF; Christiano et al., 2017) optimize for reward signals that are proxies for the desired behavior, so models may learn to exploit the reward signal rather than achieve the intended objective. For example, models reward-hack on coding tasks (Von Arx et al., 2025; MacDiarmid et al., 2025; Anthropic, 2026), generate persuasive but false content (Wen et al., 2025), exhibit sycophancy by agreeing with users instead of correcting them (Sharma et al., 2024), and fabricate actions or engage in strategic deception when these behaviors help them achieve their objectives (Scheurer et al., 2024; Transluce, 2025; Anthropic, 2026) . These findings demonstrate that post-training can reinforce undesired deceptive  \n*Equal contribution 1FAR.AI. Correspondence to: Oskar J.  \nHollinsworth \u003C[oskar@far.ai](oskar@far.ai)>, Chris Cundy \u003C[cundy@far.ai](cundy@far.ai)>.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nbehaviors, creating a need for scalable oversight methods that detect and discourage them.  \nOne promising approach is to apply trusted oversight selectively, reserving expensive supervision for responses that are most likely to be deceptive. This is the principle behind Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), a detector-guided labeling protocol to make preference learning more robust to deception. In SOLiD, the developer trains a lie detector on internal activations using a small set of ground-truth deception labels and then applies that detector during preference-data labeling to flag potentially deceptive responses. Unflagged responses are labeled by a low-cost evaluator, while flagged ones are escalated to a more expensive evaluator that more reliably selects the truthful response. Each evaluator scores both responses, and these scores are used to assign (chosen, rejected) preferences for training a reward model. In this way, SOLiD concentrates trusted supervision on the subset of responses where deception is most likely.  \nCundy & Gleave (2025) demonstrated promising results and identified a high detector true positive rate (TPR) and controlled divergence from the reference model as essential to the protocol’s success.  \nIn this paper, we test SOLiD in more realistic settings. First, scalability: we test the core protocol on Llama-3 models up to 405B parameters and Qwen-3 models up to 32B, characterizing its scaling trends. Second, practicality: we adapt SOLiD to more realistic settings. We test on-policy data, cross-dataset transfer, and a less expensive protocol variant.  \nThe core SOLiD protocol scales favorably: undetected deception decreases ","cbCaia3k4jZvvvCN","https://ap.wps.com/l/cbCaia3k4jZvvvCN","pdf",2307398,5,1,77,"English","en",105,"# Introduction\n## Scalable Oversight via Lie Detectors (SOLiD)\n# Related Work","[{\"question\":\"What problem does SOLiD address in preference learning with LLMs?\",\"answer\":\"SOLiD addresses the cost of monitoring and preventing deceptive behavior that can be reinforced by reward-based post-training like RLHF.\"},{\"question\":\"How does SOLiD scaling affect undetected deception?\",\"answer\":\"Undetected deception decreases as model size increases, with results reported as a drop from 34% for 1B models to 14% for 405B models at a 99% detector true positive rate.\"},{\"question\":\"What limitation does the paper identify for lie-detector-based oversight?\",\"answer\":\"SOLiD is sensitive to distribution shift between detector training and the preference-training data, which can drive detector false positive rates to impractical levels.\"}]","Scaling Trends for Lie Detector Oversight in Preference Learning | PDF",1784176405,194,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"scaling-trends-for-lie-detector-oversight-in-preference-learning","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/scaling-trends-for-lie-detector-oversight-in-preference-learning/81826/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does SOLiD address in preference learning with LLMs?","Question",{"text":77,"@type":78},"SOLiD addresses the cost of monitoring and preventing deceptive behavior that can be reinforced by reward-based post-training like RLHF.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does SOLiD scaling affect undetected deception?",{"text":82,"@type":78},"Undetected deception decreases as model size increases, with results reported as a drop from 34% for 1B models to 14% for 405B models at a 99% detector true positive rate.",{"name":84,"@type":75,"acceptedAnswer":85},"What limitation does the paper identify for lie-detector-based oversight?",{"text":86,"@type":78},"SOLiD is sensitive to distribution shift between detector training and the preference-training data, which can drive detector false positive rates to impractical levels.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]