[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83485-en":3,"doc-seo-83485-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83485,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","A Mechanistic View of Authority Hierarchy in LLM Sycophancy","Authority bias in language models creates safety risk by prioritizing social cues from authoritative figures over factual consistency. The study mechanistically investigates this behavior in a controlled medical QA setting using persona-specific hints that imply incorrect answers. Across Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B, responses follow a graded authority hierarchy that is not directly prompted but emerges from training. Logit lens and probing localize the effect to a critical late layer where correct representations are actively erased, scaling with authority level and only partially recovered by chain-of-thought.","A Mechanistic View of Authority Hierarchy in LLM Sycophancy  \nEmil Joswin * 1 Srujananjali Medicherla * 1 Priyanka Mary Mammen * 2  \narXiv :2607 .004 15v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nAuthority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence. We mechanistically investigate this phenomenon using a controlled medical QA setting, where hints suggesting incorrect answers are attributed to personas of varying expertise. Across Llama-3.1- 8B, Qwen3-8B, and Gemma-2-9B, we find that models respond in a graded manner proportional to perceived authority, a hierarchy that is never explicitly prompted but emerges from training.  \nLogit lens analysis and linear/non-linear probing localize this effect to a critical late layer where correct answer representations are actively erased, an erasure that scales with authority level, resists mean vector intervention, and is only partially reversible through chain-of-thought reasoning. Our findings suggest that authority-induced sycophancy is not a surface-level output bias but mechanistic knowledge erasure, a precise, layerlocalized overwriting of correct internal representations by high-status authority signals.  \n1. Introduction  \nLarge Language models are getting popular, and they have shown usefulness in a wide range of domains. Recently, even small language models with less than 10 billion parameters have shown good performance in complex reasoning tasks (Cai et al., 2025 ; Grand et al., 2025) . However, the reasoning capabilities of a model is susceptible to external cues or social context (Sharma et al., 2024) .  \nModels can learn implicit bias, just like humans, in their  \n*Equal contribution 1Independent Research 2University of Massachusetts Amherst. Correspondence to: Emil Joswin \u003C[ejoswin@gmail.com](ejoswin@gmail.com) >.  \nMechanistic Interpretability Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026 . Copyright 2026 by the author(s) .  \nCode available at [https://anonymous.4open](https://anonymous.4open) . science/r/authority-bias-llms-56C7  \ndecision-making when presented with opinions from people or sources of varying degrees of authority (Zhao et al., 2025) . This can promote sycophantic behavior where a model tries to align with users opinion rather than following correctness or logical consistency as observed in reward hacking (Perez et al., 2023) . Such behavior can be detrimental, especially when Large Language Models (LLMs) are used in critical domains such as healthcare, where we want reliable and robust answers. State-of-the-art work focuses on mitigating this behavior through various interventions, including post-training (Wei et al., 2023 ; Beigi et al., 2025), unlearning (Xing et al., 2024 ; Fang et al., 2026), and mechanistic interventions such as activation steering(Chen et al., 2025) and prompt-based strategies (Dubois et al., 2026) .  \nIn this paper, we investigate the effects of expertise levels of authority on model components when presented with a question followed by a hint from an expert persona. Specifically, we ask: does perceived authority merely bias model outputs, or does it alter internal representations in a mechanistically precise way?  \nOur contributions are as follows:  \n• We demonstrate that models respond to authority hints in a graded manner proportional to perceived expertise, a hierarchy never explicitly prompted but internalized during training (RQ1) .  \n• We localize the authority override to a critical late layer via logit lens analysis and linear/non-linear probing, identifying a sharp phase transition where correct answer representations are actively erased and overtaken by the hinted answer with erasure severity scaling with authority level (RQ2, RQ3)  \n• We show that the authority signal is question-specific and not globall","cbCaikSlHGKSYq6x","https://ap.wps.com/l/cbCaikSlHGKSYq6x","pdf",5850867,4,1,15,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Experimental Setup","[{\"question\":\"What safety concern does authority bias create in language models?\",\"answer\":\"Models can systematically prioritize authority cues over factual consistency, allowing answers to shift based on perceived source credibility rather than evidence.\"},{\"question\":\"How is authority hierarchy tested in the study?\",\"answer\":\"A controlled medical QA setting provides question-then-hint prompts where hints attributed to personas of different expertise suggest incorrect answers, enabling measurement of graded behavior.\"},{\"question\":\"What mechanistic explanation does the paper propose for authority-induced sycophancy?\",\"answer\":\"It attributes the effect to mechanistic erasure of correct internal representations in a critical late layer, with erasure strength scaling with authority level.\"}]",1784188347,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-mechanistic-view-of-authority-hierarchy-in-llm-sycophancy","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/a-mechanistic-view-of-authority-hierarchy-in-llm-sycophancy/83485/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What safety concern does authority bias create in language models?","Question",{"text":75,"@type":76},"Models can systematically prioritize authority cues over factual consistency, allowing answers to shift based on perceived source credibility rather than evidence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is authority hierarchy tested in the study?",{"text":80,"@type":76},"A controlled medical QA setting provides question-then-hint prompts where hints attributed to personas of different expertise suggest incorrect answers, enabling measurement of graded behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanistic explanation does the paper propose for authority-induced sycophancy?",{"text":84,"@type":76},"It attributes the effect to mechanistic erasure of correct internal representations in a critical late layer, with erasure strength scaling with authority level.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]