[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84519-en":3,"doc-seo-84519-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84519,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training","Standard post-training pipelines—supervised fine-tuning (SFT) and reinforcement learning (RL)—aim to make language models helpful, yet they can weaken values learned earlier. This work tests whether the domain of post-training data changes that effect. A Llama 3.1 8B model mid-trained on compassion-oriented synthetic data is post-trained on helpfulness or coding; helpfulness degrades animal compassion far more, and it also reduces general moral reasoning on English items. Coding post-training better preserves mid-trained values across languages without measurable harm to general reasoning.","arXiv :2606 .26 102v 3 [ cs .CL] 12 Jul 2026  \nHelpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training  \nJuliana Seawell∗† Jasmine Brazilek∗† Miles Tidmarsh††Compassion Aligned Machine Learning (CaML) ∗Equal contribution  \nAbstract  \nStandard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes can degrade values instilled earlier in training. We ask whether the domain of the post-training data matters.  \nStarting from a Llama 3.1 8B model mid-trained on compassion-oriented synthetic data, we post-train on helpfulness data (Dolly-15k for SFT, RLHFlow for GRPO) or on coding data (Magicoder for both), and evaluate on ANIMA 2.2 (Animal Norms In Moral Assessment) 1 and on MORU (Moral Reasoning Under Uncertainty) . Helpfulness training degrades animal compassion far more than coding training (ANIMA, SFT: 35.7% vs. 65.2%; GRPO: 15.4% vs. 30.3%), and the gap replicates across two independent helpfulness datasets and two training paradigms. On English MORU items, helpfulness training also degrades general moral reasoning, by 25.5 percentage points (46.4% vs.  \n71.9%), but this gap disappears on the multilingual MORU set (52.3% vs. 51.2%) . The compassion effect, in contrast, transfers across languages: coding post-training beats helpfulness post-training by 32 percentage points on English ANIMA items and by  \n26 points on non-English items (both p \u003C 10 −6), even though all post-training data was English. Values instilled through mid-training appear to be encoded more deeply, and more cross-lingually, than reasoning changes from domain-specific post-training. For labs building on value-laden mid-training, coding-domain post-training preserves mid-trained values much better than helpfulness post-training, at no measured cost to general reasoning.  \n1 Introduction  \nLarge language models are typically aligned through a pipeline of pre-training followed by supervised fine-tuning (SFT) and reinforcement learning (RL) . The dominant alignment paradigm targets the HHH criteria: Helpful, Honest, and Harmless (Askell et al., 2021), but these objectives can conflict. Bai et al. (2022) demonstrated quantitatively that helpfulness and harmlessness stand in trade-off: preference models trained primarily on one quality degrade the other. Lin et al. (2024) showed that pushing a model to be more aligned can measurably degrade its core capabilities, such as reading comprehension. Recent theoretical work suggests that the cost of alignment scales with how much the safety-relevant direction in weight space overlaps with the capability subspace (Chen et al., 2025a); the greater the overlap, the costlier alignment becomes.  \nThe fragility of aligned values under fine-tuning is by now established empirically. Qi et al. (2023) showed that safety alignment can be compromised by as few as 10 adversarially designed examples, and that even benign fine-tuning on commonly used datasets such as Alpaca and Dolly inadvertently  \n1Previously released as the Animal Harm Benchmark (AHB); renamed in May 2026 to disambiguate from an unrelated benchmark another group released under the AHB name. The questions, dimensions, and scoring are unchanged.  \ndegrades safety. The Superficial Alignment Hypothesis suggests that alignment primarily modifies the model’s output format rather than its deep knowledge, implying that values are truly encoded during pre-training and are vulnerable to displacement when SFT shifts the model’s output distribution. Chen et al. (2025b) have recently shown that model traits like sycophancy and hallucination can be tracked as specific directions in the model’s internal representations (persona vectors), and that fine-tuning on different domains shifts these traits in predictable ways, with larger shifts when the fine-tuning data is more similar to the trait being measured.  \nA key finding for the present work is t","cbCaifSqD9K2FHxK","https://ap.wps.com/l/cbCaifSqD9K2FHxK","pdf",275298,1,23,"English","en",105,"# Abstract\n# Introduction\n## Alignment pipeline and value degradation\n## Domain effects on SFT/RL\n## Non-anthropocentric values and RL feedback\n## Harmlessness evaluation toward non-human entities","[{\"question\":\"What does the document investigate about post-training?\",\"answer\":\"It examines whether the domain of post-training data (helpfulness versus coding) affects how values learned during mid-training degrade.\"},{\"question\":\"How does helpfulness post-training impact compassion compared with coding post-training?\",\"answer\":\"Helpfulness training degrades animal compassion much more than coding training across multiple datasets and training paradigms.\"},{\"question\":\"Does domain-specific post-training also affect moral reasoning beyond compassion?\",\"answer\":\"On English MORU items, helpfulness post-training reduces general moral reasoning, while the gap disappears on the multilingual MORU set. Coding-domain post-training better preserves mid-trained values across languages.\"}]",1784196273,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"helpfulness-hurts-domain-dependent-degradation-of-mid-trained-compassion-values-under-post-training","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/helpfulness-hurts-domain-dependent-degradation-of-mid-trained-compassion-values-under-post-training/84519/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the document investigate about post-training?","Question",{"text":75,"@type":76},"It examines whether the domain of post-training data (helpfulness versus coding) affects how values learned during mid-training degrade.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does helpfulness post-training impact compassion compared with coding post-training?",{"text":80,"@type":76},"Helpfulness training degrades animal compassion much more than coding training across multiple datasets and training paradigms.",{"name":82,"@type":73,"acceptedAnswer":83},"Does domain-specific post-training also affect moral reasoning beyond compassion?",{"text":84,"@type":76},"On English MORU items, helpfulness post-training reduces general moral reasoning, while the gap disappears on the multilingual MORU set. Coding-domain post-training better preserves mid-trained values across languages.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]