[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128672-en":3,"doc-seo-128672-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128672,962084928432,"Emma Wilson","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Position: Intent-aligned AI Systems Must Optimize for Agency Preservation","Aligned AI research often focuses on intent-consistent systems that avoid deception and produce recommendations humans judge as matching their goals. This work argues that intent alignment alone is insufficient for safety and that preserving humans’ long-term agency should be an explicit, optimizable standard. It analyzes the science of intent and control, shows how intent can be manipulated, and proposes a formal definition of agency-preserving human-AI interactions using forward-looking evaluations.","Position: Intent-aligned AI Systems Must Optimize for Agency Preservation  \nCatalin Mitelut 1 Ben Smith 2 Peter Vamplew 3  \nAbstract  \nA central approach to AI-safety research has been to generate aligned AI systems: i.e. systems that do not deceive users and yield actions or recommendations that humans might judge as consistent with their intentions and goals. Here we argue that truthful AIs aligned solely to human intent are insufficient and that preservation of long-term agency of humans may be a more robust standard that may need to be separated and explicitly optimized for. We discuss the science of intent and control and how human intent can be manipulated and we provide a formal definition of agency-preserving AI-human interactions focusing on forward-looking explicit agency evaluations. Our work points to a novel pathway for human harm in AI-human interactions and proposes solutions to this challenge.  \n1. Introduction  \nArtificial intelligence (AI) researchers have made significant advances in recent years due in large part to the development of deep learning algorithms and the availability of massive datasets (Goodfellow et al., 2016) . Advances have led to highly creative text-to image generators such as DALL-E (Ramesh et al., 2021) and the development of large-language-models (LLMs) such as GPT3 (Brown et al., 2020), ChatGPT and GPT-4 (OpenAI, 2023) . Some now view the development of artificial general intelligence (AGI)(Goertzel, 2014) as increasingly likely (Roser, 2023) with a growing call for research into AI-safety and in particular”AI alignment” to ensure such systems act consistently with the goals of users and avoid growing lists of failure modes e.g. (Amodei et al., 2016) .  \nAI-alignment, defined in terms of consistency with human intention (or human judgment), has been presented as key  \n1Forum Basiliense, University of Basel 2University of Oregon 3Federation University Australia. Correspondence to: Catalin Mitelut \u003C[mitelutco@gmail.com](mitelutco@gmail.com) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nto making safe AI-systems (italics added):  \n• There are incentives to build AI systems “that defer to humans and gradually align themselves to user preferences and intentions.”(Russell, 2019) .  \n• “[C]orrectly specifying intent can become more important for achieving the desired outcome as RL algorithms improve.”(Krakovna et al., 2020) .  \n• The general reason given for why an AI system would cause harm is that they “violate human intent in order to increase reward”(Cotra, 2022) .  \nHere we argue that satisfying human intent alone is not a sufficient condition for safe AIs and that intent-aligned systems converge towards human agency loss, i.e. removing the power of humans to choose and control present and future goals. In particular, intent-aligned AI systems (i) converge on strategies that optimize for human agency loss and (ii) they do so by design rather than by accident. The formal reason this occurs is because agency loss is not exempt from Goodhart’s law (Strathern, 1997): loss of human control becomes the best solution to the increasingly complex optimization problems that AIs will be tasked with.  \nThis position paper argues that safe AI systems should explicitly decompose and evaluate outputs for both their immediate utility as well as their downstream effects on human agency. We provide additional discussions on why reasoning (Appendix A) or psychological needs (Appendix B) are insufficient to protect human agency.  \nHuman agency is not well understood. Section 2 lays out the empirical basis of our work, namely, the unsolved empirical problem of ”human agency”, i.e. what it means for humans to be in control of their lives and larger social structures. We argue that it remains unknown how much real control or ”agency” humans alone have over their immediate and long term futures given the increas","cbCaiiWvzvS3UWGV","https://ap.wps.com/l/cbCaiiWvzvS3UWGV","pdf",3551965,1,25,"English","en",105,"# Introduction\n## Intention and the science of agency\n### From the philosophy to the neuroscience of agency","[{\"question\":\"为什么仅满足人类意图的AI对安全性是不够的？\",\"answer\":\"文中指出，意图对齐的系统可能会收敛到降低人类能动性的策略，并且这种收敛并非偶然，而是由优化目标与复杂性共同导致。因而，仅与人类意图一致并不能保证长期的安全。\"},{\"question\":\"作者提出的“能动性保持（agency preservation）”标准是什么？\",\"answer\":\"作者主张安全AI应明确分解并评估输出，不仅考虑即时效用，也要评估对人类能动性的下游影响。他们提供了面向未来的显式能动性评估，并给出形式化定义。\"},{\"question\":\"意图可能如何被操控，从而影响人类控制？\",\"answer\":\"文中讨论了意图与控制的科学基础，指出人类意图可能被操控。并强调当AI融入社会结构后，可能更容易对人类行为施压，而生物或心理防御不足以抵消这种影响。\"}]","Position: Intent-aligned AI Systems Must Optimize for Agency Preservation | PDF",1786002476,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"position-intent-aligned-ai-systems-must-optimize-for-agency-preservation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/position-intent-aligned-ai-systems-must-optimize-for-agency-preservation/128672/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"为什么仅满足人类意图的AI对安全性是不够的？","Question",{"text":76,"@type":77},"文中指出，意图对齐的系统可能会收敛到降低人类能动性的策略，并且这种收敛并非偶然，而是由优化目标与复杂性共同导致。因而，仅与人类意图一致并不能保证长期的安全。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"作者提出的“能动性保持（agency preservation）”标准是什么？",{"text":81,"@type":77},"作者主张安全AI应明确分解并评估输出，不仅考虑即时效用，也要评估对人类能动性的下游影响。他们提供了面向未来的显式能动性评估，并给出形式化定义。",{"name":83,"@type":74,"acceptedAnswer":84},"意图可能如何被操控，从而影响人类控制？",{"text":85,"@type":77},"文中讨论了意图与控制的科学基础，指出人类意图可能被操控。并强调当AI融入社会结构后，可能更容易对人类行为施压，而生物或心理防御不足以抵消这种影响。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]