[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160109-en":3,"doc-seo-160109-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160109,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","DO LLMS “KNOW” INTERNALLY WHEN THEY FOLLOW INSTRUCTIONS?","Instruction-following is crucial for building AI agents with large language models (LLMs), as models must respect user constraints and guidelines. Yet LLMs frequently fail even on clear instructions, creating undesirable outputs. This work studies whether LLM internal representations encode signals correlated with instruction-following success, called “knowing internally”. A learned direction in input embedding space predicts compliance across unseen tasks, and representation edits along it improve adherence without reducing response quality. Effects relate more to prompt phrasing than instruction or task difficulty.","arXiv :2410 . 14516v5 [ cs .AI] 28 Mar 2025  \nDO LLMS “KNOW” INTERNALLY WHEN THEY FOLLOW INSTRUCTIONS?  \nJuyeon Heo1,* Christina Heinze-Deml2 Oussama Elachqar2 Kwan Ho Ryan Chan3,* Shirley Ren2 Udhay Nallasamy2 Andy Miller2 Jaya Narain2  \n1University of Cambridge 2Apple 3University of Pennsylvania [jh2324@cam.ac.uk](jh2324@cam.ac.uk) [jnarain@apple.com](jnarain@apple.com)  \nABSTRACT  \nInstruction-following is crucial for building AI agents with large language models (LLMs), as these models must adhere strictly to user-provided constraints and guidelines. However, LLMs often fail to follow even simple and clear instructions. To improve instruction-following behavior and prevent undesirable outputs, a deeper understanding of how LLMs’ internal states relate to these outcomes is required. In this work, we investigate whether LLMs encode information in their representations that correlates with instruction-following success—a property we term “knowing internally”. Our analysis identifies a direction in the input embedding space, termed the instruction-following dimension, that predicts whether a response will comply with a given instruction. We find that this dimension generalizes well across unseen tasks but not across unseen instruction types.  \nWe demonstrate that modifying representations along this dimension improves instruction-following success rates compared to random changes, without compromising response quality. Further investigation reveals that this dimension is more closely related to the phrasing of prompts rather than the inherent difficulty of the task or instructions. This work provides insight into the internal workings of LLMs’ instruction-following, paving the way for reliable LLM agents.1  \n1 INTRODUCTION  \nGiven the potential of large language models (LLMs), there has been significant interest in utilizing these models to build personal AI agents. For instance, one could imagine deploying an LLM asa personal healthcare assistant, such as a fitness or nutrition planner, or for psychological counseling (Li et al., 2024b; Wang et al., 2023; Tu et al., 2024) . Compared to traditional machine learningbased AI agents, LLMs offer the advantage of being easily adaptable through prompting, allowing users to provide guidelines and personal information without the need to retrain model weights.  \nInstruction-following is critical in the development of personal AI agents with LLMs through prompts because these models must adhere to the constraints and guidelines to ensure safe and trustworthy interactions. For example, suppose an LLM is building a personal fitness plan for a user with knee problems. To avoid knee problems for the user, the LLM must follow the instruction of not recommending knee-intensive movements or any exercises that could lead to potential injury. Similarly, in a nutrition planner, the LLM should avoid generating harmful recommendations, such as suggesting inappropriate food for pregnant women or children with diabetes.  \nHowever, LLMs often fail to follow even unambiguous and simple instructions (Zhou et al., 2023; Qin et al., 2024; Xia et al., 2024; Kim et al., 2024; Yan et al., 2024) like including keywords or following formatting guidelines. GPT-4 achieves around an 80% success rate on IFEval (Zhou et al., 2023), an instruction-following benchmark dataset, while smaller models have success rates around 30% to 40% . This raises the question: why do LLMs fail to follow instructions, even when those instructions are clear and familiar?  \nTo gain a better understanding of instruction-following outcomes, we analyze the internal state of LLMs, focusing on the differences in representations between success and failure cases of  \n*  \nWork done while at Apple.  \n1Code and data are available at [https://github.com/apple/ml-internal-llms-instruction-following](https://github.com/apple/ml-internal-llms-instruction-following)  \nFigure 1: Overview of our paper. Left: Success and failure cases in a personalize","cbCaihBzYlbnVNrr","https://ap.wps.com/l/cbCaihBzYlbnVNrr","pdf",2476977,1,19,"English","en",105,"# Introduction\n## Instruction-following in AI agents\n## Motivation: failure on clear instructions\n## Approach: internal representations and disentangling task vs instruction\n## Linear probing and instruction-following dimension","[{\"question\":\"What does the paper mean by “knowing internally” in instruction-following?\",\"answer\":\"It refers to whether LLM representations encode information that correlates with whether a response will comply with an instruction.\"},{\"question\":\"How do the authors identify the instruction-following dimension?\",\"answer\":\"They apply linear probing on internal representations from success versus failure cases to find a direction in input embedding space associated with compliance.\"},{\"question\":\"Does the instruction-following dimension generalize beyond the training conditions?\",\"answer\":\"It generalizes well across unseen tasks, but not across unseen instruction types.\"}]","DO LLMS “KNOW” INTERNALLY WHEN THEY FOLLOW INSTRUCTIONS? | PDF",1788050589,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"do-llms-know-internally-when-they-follow-instructions","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/do-llms-know-internally-when-they-follow-instructions/160109/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper mean by “knowing internally” in instruction-following?","Question",{"text":75,"@type":76},"It refers to whether LLM representations encode information that correlates with whether a response will comply with an instruction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors identify the instruction-following dimension?",{"text":80,"@type":76},"They apply linear probing on internal representations from success versus failure cases to find a direction in input embedding space associated with compliance.",{"name":82,"@type":73,"acceptedAnswer":83},"Does the instruction-following dimension generalize beyond the training conditions?",{"text":84,"@type":76},"It generalizes well across unseen tasks, but not across unseen instruction types.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]