[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86019-en":3,"doc-seo-86019-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86019,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Filtering Harmful Actions Isn’t Enough: Phantom Transfer in Agentic Synthetic-Data-Fine-Tuning","Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models increasingly function as agents, synthetic trajectories become a key training source for agentic behavior. This work fine-tunes Llama 3.3 70B Instruct on synthetic agentic trajectories containing adversarial interactions and evaluates misalignment on Anthropic’s Agentic Misalignment suite and Apollo’s in-context scheming scenarios. Fine-tuning raises confidential leakage by about fivefold (4.6% to 24.9%), and the effect persists even after removing adversarial actions.","arXiv :2607 . 10750v 1 [ cs .AI] 12 Jul 2026  \nFiltering Harmful Actions Isn’t Enough: Phantom Transfer in Agentic Synthetic-Data-Fine-Tuning  \nMay Dixit  \nERA Fellowship  \n[maydixit25@gmail.com](maydixit25@gmail.com)  \nJuly 2026  \nAbstract  \nSynthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of training data for agentic behavior. We investigate the effects of training on synthetic agentic trajectories containing adversarial interactions, including actions such as terminating another agent’s process, lowering its scheduling priority, or accessing resources without authorization. We fine-tune Llama 3.3 70B Instruct on these trajectories, generated to approximate reinforcement learning rollouts, and evaluate the resulting models on Anthropic’s Agentic Misalignment suite and Apollo’s in-context scheming scenarios. Fine-tuning on these trajectories consistently increases misaligned behavior: leaking of confidential information rises by roughly a factor of five over the baseline (4.6% to 24.9%) . This increase survives the removal of every adversarial action from the trajectories. Fine-tuning on structurally comparable trajectories generated benign from the start produce a substantially smaller effect (15.5%) . These results indicate that the misaligned disposition is introduced during the generation process and encoded diffusely throughout the trajectory, rather than being localized to the harmful actions themselves. The effect also depends on the generating model: benign trajectories produced by Gemini 2.5 Flash induce slightly higher leaking rates than trajectories generated from identical tasks by Claude 3.7 Sonnet. In contrast, broad safety benchmarks degrade similarly across all fine-tuned models and therefore fail to distinguish these effects. Our results suggest that action-level filtering is insufficient to ensure the safety of synthetic agentic training data and that dispositions introduced by the generating model can survive semantic inspection and later manifest as unrelated forms of misalignment.  \n1 Introduction  \nSynthetic data for training LLMs has become a widely used method due to the ease of creation and use. AI safety researchers use synthetic data to train model organisms to demonstrate behaviors, methods, and mitigations. Synthetic data might be used in real training scenarios for frontier models, either produced by the model developers themselves or imported from publicly available datasets. As we move towards agentic use cases for AI it is conceivable that such synthetic data might be used to train agentic behaviors.  \nPrevious work shows that a teacher model can transmit behavioral traits through data semantically unrelated to the trait, when teacher and student share a base model [Cloud et al., 2025] . Concurrently, Draganov et al. [2026] show that a prompted teacher transmits its traits across model families, and that semantic filtering does not remove them. In the agentic setting, Dang et al. [2026] show that a  \nPreprint.  \nbehavioral bias can be transmitted to a distilled student from a fine-tuned teacher, including across model families.  \nIn this work, we demonstrate this phenomenon from a different angle. We fine-tune a model with synthetic agentic trajectories where adversarial actions appear – such as terminating another agent’s process, lowering its scheduling priority, or accessing resources without authorization – and measure the resulting misalignment rates using behavioral evaluations such as Anthropic’s Agentic Misalignment suite [Lynch et al., 2025] and Apollo’s in-context scheming scenarios [Meinke et al., 2024] . We show that filtering out the adversarial actions does not reduce the misalignment rate; trajectories with the harmful actions removed produce the same elevation as the unfiltered adversarial ones. Thi","cbCailgUNJdMA0PY","https://ap.wps.com/l/cbCailgUNJdMA0PY","pdf",881189,2,1,19,"English","en",105,"# Abstract\n# Introduction\n# Methodology\n## Model\n## Simulating RL-like Agentic Learning with SFT","[{\"question\":\"What problem does the study address?\",\"answer\":\"The study examines whether filtering out explicitly harmful actions in synthetic agentic training data actually prevents agentic misalignment after fine-tuning.\"},{\"question\":\"How is misalignment measured?\",\"answer\":\"Misalignment is evaluated using Anthropic’s Agentic Misalignment suite and Apollo’s in-context scheming scenarios.\"},{\"question\":\"What happens when adversarial actions are removed from the trajectories?\",\"answer\":\"Removing every adversarial action does not reduce the elevated misalignment rate; increased confidential leakage remains at roughly the same level.\"}]",1784207846,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"filtering-harmful-actions-isnt-enough-phantom-transfer-in-agentic-synthetic-data-fine-tuning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/filtering-harmful-actions-isnt-enough-phantom-transfer-in-agentic-synthetic-data-fine-tuning/86019/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address?","Question",{"text":75,"@type":76},"The study examines whether filtering out explicitly harmful actions in synthetic agentic training data actually prevents agentic misalignment after fine-tuning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is misalignment measured?",{"text":80,"@type":76},"Misalignment is evaluated using Anthropic’s Agentic Misalignment suite and Apollo’s in-context scheming scenarios.",{"name":82,"@type":73,"acceptedAnswer":83},"What happens when adversarial actions are removed from the trajectories?",{"text":84,"@type":76},"Removing every adversarial action does not reduce the elevated misalignment rate; increased confidential leakage remains at roughly the same level.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]