[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82460-en":3,"doc-seo-82460-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82460,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust","Language models have evolved from unreliable text generators into large models with trillions of parameters, yet their internal representations remain difficult to interpret and govern. With growing reliance on LLMs for tool use and medium-to-high-stakes decisions, the document proposes methods to control model behavior and to decide when to trust outputs. It introduces steering vectors for control and latent-space model calibrators for trustworthiness, clarifying latent space use for safer language technology.","Harnessing the Latent Space:  \nFrom Steering Vectors to Model Calibrators for Control and Trust  \nNishant Subramani   \n: Carnegie Mellon University, Language Technologies Institute [nishant2@cs.cmu.edu](nishant2@cs.cmu.edu)  \narXiv :2607 .00083v1 [ cs .CL] 30 Jun 2026  \nAbstract  \nLanguage models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increases come hand-in-hand with increases in scale, making understanding the internal representations of models more challenging. Since millions of users increasing rely on language models to interact with external tools or make decisions in medium or high-stakes scenarios, we need to establish control over model behavior and know when to trust model outputs. In this paper, we discuss our contributions on harnessing the latent spaces by proposing steering vectors for control and developing latent spacebased model calibrators for trust. Together, our contributions help demystify the latent spaces of language models and offer new insights into how to harness model internals to build more trustworthy language technology.  \n1 Introduction  \nNeural network language models (LMs) have evolved from small, unreliable text generators to very large models capable of solving complex reasoning tasks (Peters et al., 2018 ; Radford et al., 2019 ; Groeneveld et al., 2024 ; Yang et al., 2025 ; Team et al., 2025, inter alia) . Despite the vast capability increases, analyzing the internal representations of trillion-parameter models is challenging. Due to this, the NLP community has increasingly treated models as black boxes, neglecting understanding the inner-workings of models. Even though large language models (LLMs) are scaled to millions of users, increasingly interact with external tools (Qu et al., 2024), and make decisions in medium and high-stakes scenarios (Thirunavukarasu et al., 2023), we rely by-andlarge on simple behavioral observation (Hendryckset al., 2021 ; Srivastava et al., 2023 ; Liang et al., 2023, inter alia) . As a community, we must build fundamental understanding of the inner-workings  \nFigure 1: Our contributions on harnessing the latent spaces of language models: §2 and §3 focus on control, proposing steering vectors for the first time for LSTMsand transformer-based models. §4 and §5 focus on trust, building model-internal confidence estimators to assess confidence of language model output generations.  \nof models and operationalize the internal representations of LLMs. We need to establish control over model behavior to ensure safety and alignment and establish confidence estimation mechanisms which can accurately adjudicate trust.  \nWe present four threads of research aimed to demystify and harness the latent spaces of language models. To achieve model control, we show that LSTM-based language models can be minimally steered for exact generation (§2) . We then adapt to transformer-based models in §3, showing both fine-grained and coarse-grained control via exact and concept-based steering. Shifting to trustworthiness, we build model-internal confidence estimators (MICE) to calibrate LLM generations in tool-calling scenarios (§4) . Lastly, we broaden the framework to new model families and tasks by proposing activation-based confidence, utility, and trust estimators (ACUTE; §5) . Together, our contributions offer actionable recipes to harness model internals to build more controllable, well-calibrated, and trustworthy language technologies.  \nFigure 2: Here we show how the steering vector zsteer can be injected into an LSTM-based language model (left) and a transformer-based one (right) . On the left, Z is shown to have a larger dimension than the model dimension. If dim(Z) equals the model dimension, K = 1, and thus there is just one vector z 1.  \n2 Control: Steering Vectors for LSTMs (Subramani et al., 2019)  \nWe focus on control, specifically trying to answer one key question:  \nKey Question 1  \nCan LSTM-ba","cbCaicI7EKGPlm9G","https://ap.wps.com/l/cbCaicI7EKGPlm9G","pdf",4922662,3,1,12,"English","en",105,"# Abstract\n# Introduction\n# Control: Steering Vectors for LSTMs\n## Prior Work\n## Our Contributions","[{\"question\":\"为什么需要对语言模型进行控制与“何时可信”的判断？\",\"answer\":\"随着语言模型被更多用户用于工具交互以及中高风险决策，仅靠简单的行为观察不足以保证安全与一致性，因此需要控制模型行为并建立能够准确裁决可信度的机制。\"},{\"question\":\"文中提出的“控制”方法核心是什么？\",\"answer\":\"面向控制问题，提出使用 steering vectors 来实现对语言模型输出的精确生成（在不更新单个参数的前提下），并将该思路扩展到 transformer 等模型，支持精细与粗粒度控制。\"},{\"question\":\"文中如何实现对模型输出的“信任校准”？\",\"answer\":\"提出模型内部置信度估计器（MICE），在工具调用场景中对 LLM 生成进行校准；进一步通过激活相关的置信度、效用与信任估计器（ACUTE）扩展到新的模型家族与任务。\"}]",1784180596,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"harnessing-the-latent-space-from-steering-vectors-to-model-calibrators-for-control-and-trust","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/harnessing-the-latent-space-from-steering-vectors-to-model-calibrators-for-control-and-trust/82460/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么需要对语言模型进行控制与“何时可信”的判断？","Question",{"text":75,"@type":76},"随着语言模型被更多用户用于工具交互以及中高风险决策，仅靠简单的行为观察不足以保证安全与一致性，因此需要控制模型行为并建立能够准确裁决可信度的机制。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文中提出的“控制”方法核心是什么？",{"text":80,"@type":76},"面向控制问题，提出使用 steering vectors 来实现对语言模型输出的精确生成（在不更新单个参数的前提下），并将该思路扩展到 transformer 等模型，支持精细与粗粒度控制。",{"name":82,"@type":73,"acceptedAnswer":83},"文中如何实现对模型输出的“信任校准”？",{"text":84,"@type":76},"提出模型内部置信度估计器（MICE），在工具调用场景中对 LLM 生成进行校准；进一步通过激活相关的置信度、效用与信任估计器（ACUTE）扩展到新的模型家族与任务。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]