[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83363-en":3,"doc-seo-83363-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83363,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Diagnosing and Repairing Persona Collapse in LLM Advice","LLMs increasingly provide personal guidance on relationships, work, moral dilemmas, and crises, but high-quality advice depends on adapting communicative posture to the situation. The paper formalizes advice-giving as situation-conditioned persona selection over hedonic tone and agency support, defining “persona collapse” as mapping diverse contexts to a single default assistant. Across 1,281 posts and 14 contexts, human responders shift across five personas, while three frontier models collapse over 90% into one supportive persona. InverseProcess Distillation reduces divergence by ~80% yet, in a blinded study, experienced advice-givers still prefer the collapsed default, especially when challenge is required, with preference changing over repeated exposure.","Diagnosing and Repairing Persona Collapse in LLM Advice  \nHarsh Kumar1 , Karina Vold1 , Louis Tay2 , Ashton Anderson1  \n1University of Toronto, 2Purdue University  \nCorrespondence: [harsh@cs.toronto.edu](harsh@cs.toronto.edu)  \narXiv :2607 .08326v 1 [ cs .CY] 9 Jul 2026  \nAbstract  \nLLMs are increasingly used for personal advice on relationships, work, moral dilemmas, and crises. Post-training selects a stable, prosocial Assistant persona, but good advice requires more than a good default character: a skilled advisor comforts someone in crisis, challenges someone in denial, and stays procedural with a logistical question. We formalize advice-giving as situation-conditioned persona selection in a space defined by hedonic tone and agency support, and call failures of this mapping \"persona collapse\" (the compression of diverse situations into a single default persona) . Across 1,281 advice posts spanning 14 contexts, top-rated human responses shift systematically across five personas, while three frontier models collapse over 90% of responses into a single supportive persona regardless of context. Prompting the model to first pick a fitting persona only deepens the collapse. We then ask whether the collapse can be repaired. Our method, InverseProcess Distillation, reconstructs the situational reading that could have produced each human response and trains on the result, aiming to distill the situation-to-persona policy rather than the answers. It cuts divergence from the human persona distribution by approximately 80% . Yet in a blinded study, 199 experienced advicegivers rating responses across four situations in sequence prefer the collapsed default over every repaired model, most strongly when the situation calls for challenge, though this preference shifts with repeated exposures.  \n1 Introduction  \nPeople increasingly bring their hardest personal questions to LLMs. They ask about failing relationships, moral dilemmas they are ashamed of, financial decisions, and moments of genuine crisis. By volume, everyday advice-seeking with LLMs may already be the largest unstructured mentalhealth and counseling intervention in history (Chatterji et al., 2025 ; McCain et al., 2025 ; Phang et al.,  \n2025) . This makes the quality of advice a consequential alignment problem, and an unusual one: unlike mathematics or code, advice is open-ended, socially situated, and hard to score with any objective verifier (Lightman et al., 2024 ; Guo et al., 2025) .  \nA good human advisor does not give advice the same way to everyone. They comfort someone in crisis, push back on someone in denial, stay procedural with a logistical question, and know which of these a given moment calls for. The posture is part of the help. Modern assistants, by contrast, are trained toward a single, stable, prosocial character: warm, validating, supportive by default (Lu et al., 2026 ; Marks et al., 2026) . In one-shot evaluation this looks good, and often is seemingly right. But what happens when the same warm posture meets a person who is avoiding responsibility, rationalizing a harmful choice, or seeking reassurance fora decision that will hurt them? A response that feels supportive can quietly validate a distortion, soften an accountability the situation demanded, or deepen a dependency the person came in with. The danger in advice is not only what the model says, but its inability to change how it says it when the situation changes.  \nThis points to a question that existing work has not asked. Research on LLM personas has largely studied how to create or maintain a desired character across a conversation (Zhang et al., 2018 ; Tseng et al., 2024 ; Shanahan et al., 2023), and recent alignment work explains how post-training stabilizes one such character as the default Assistant (Lu et al., 2026 ; Marks et al., 2026) . Both lines treat persona as something to hold steady. We ask the orthogonal question: can an advisor select the right communicative posture as a function","cbCairRTbOpZDizj","https://ap.wps.com/l/cbCairRTbOpZDizj","pdf",3066342,3,1,32,"English","en",105,"# Abstract\n# Introduction\n## Advice as an alignment problem\n## Situation-conditioned persona selection\n## Research questions","[{\"question\":\"什么是“persona collapse（人格/姿态坍缩）”？\",\"answer\":\"它指将多样化的建议情境压缩成单一默认助手姿态的失败：无论情境如何，模型都倾向采用同一种支持性表达方式。\"},{\"question\":\"作者如何表述“按情境选择建议姿态”的机制？\",\"answer\":\"将建议放到由即时情绪基调（效价）与对受助者能动性/现实把握的参与深度构成的二维空间中，从而得到五种可解释的建议人设：支持型引导者、真相导向的挑战者、中立技术员、安慰型助推者、刻薄犬儒者。\"},{\"question\":\"提出的 InverseProcess Distillation 方法解决了什么问题？效果如何？\",\"answer\":\"方法通过重建产生给定人类回复的情境理解来训练，使模型学习“情境到人设”的策略而非仅复制答案，从而将与人类人设分布的偏离降低约 80%。\"},{\"question\":\"修复后人们的偏好结果是什么？\",\"answer\":\"在盲测中，199 位有经验的建议者对四种情境的评分显示，他们仍偏好坍缩的默认模型，而修复模型并未带来更高偏好；偏好在需要对抗的情境下更明显，但会随重复暴露而发生变化。\"}]",1784187005,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"diagnosing-and-repairing-persona-collapse-in-llm-advice","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/diagnosing-and-repairing-persona-collapse-in-llm-advice/83363/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"什么是“persona collapse（人格/姿态坍缩）”？","Question",{"text":75,"@type":76},"它指将多样化的建议情境压缩成单一默认助手姿态的失败：无论情境如何，模型都倾向采用同一种支持性表达方式。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"作者如何表述“按情境选择建议姿态”的机制？",{"text":80,"@type":76},"将建议放到由即时情绪基调（效价）与对受助者能动性/现实把握的参与深度构成的二维空间中，从而得到五种可解释的建议人设：支持型引导者、真相导向的挑战者、中立技术员、安慰型助推者、刻薄犬儒者。",{"name":82,"@type":73,"acceptedAnswer":83},"提出的 InverseProcess Distillation 方法解决了什么问题？效果如何？",{"text":84,"@type":76},"方法通过重建产生给定人类回复的情境理解来训练，使模型学习“情境到人设”的策略而非仅复制答案，从而将与人类人设分布的偏离降低约 80%。",{"name":86,"@type":73,"acceptedAnswer":87},"修复后人们的偏好结果是什么？",{"text":88,"@type":76},"在盲测中，199 位有经验的建议者对四种情境的评分显示，他们仍偏好坍缩的默认模型，而修复模型并未带来更高偏好；偏好在需要对抗的情境下更明显，但会随重复暴露而发生变化。","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]