[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85149-en":3,"doc-seo-85149-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85149,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States","The study tests whether identity-specifying system prompts create statistically distinguishable geometric fingerprints in token-indexed transformer hidden-state trajectories across four open-weight language models and post-training regimes. Using a controlled three-prompt setup and five geometry metrics on k-NN trajectory graphs, including Wasserstein distance over Ollivier-Ricci curvature edge distributions and multiple anisotropy and alignment controls, results show a qualitative migration of the identity signal from direction-coded encoding in the base model to magnitude-coded encoding in multimodal instruction-tuning, absent in other regimes.","arXiv :2607 .09842v 1 [ cs .LG] 10 Jul 2026  \nFrom Direction to Magnitude:  \nHow Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer  \nHidden States  \nJorge A. Castillo 1 , Marco Torres Yévenes 1 , and Juan Carlos Lanas 1  \n1 Axis Dynamics SpA, Santiago, Chile, [jorge. castillo@axisdynamics. cl](jorge. castillo@axisdynamics. cl), [mtorres@axisdynamics. cl](mtorres@axisdynamics. cl), [jc@axisdynamics. cl](jc@axisdynamics. cl)  \nJuly 14, 2026  \nAbstract  \nWe investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the token-indexed hidden-state trajectories of four openweight transformer language models spanning four post-training regimes: no training (Gemma- 4-E4B base), multimodal RLHF (Gemma-4-E4B-it), RL distillation (DeepSeek-R1-DistillQwen-7B), and supervised instruction-tuning (Qwen2.5-7B-Instruct) . Three controlled prompt conditions (an identity-specifying axis prompt ∼2129 tokens, a length-matched generic-assistant prompt, and a 26-token vanilla baseline) are compared via five geometric metrics with distinct theoretical anchors: the 1-Wasserstein distance between edge-wise distributions of Ollivier-Ricci curvature on k-NN trajectory graphs, the prompt-response alignment with all-but-the-top anisotropy correction, the initial-state cosine, the PCA-50 silhouette of axis-vs-generic clustering, and the inter-trajectory cosine consistency. All inferential claims are based on trajectory-level permutation null distributions and on multiple geometric controls (teacher-forced content controls, temporal-chain versus k-NN graph topology, ABT-projected k-NN, angular versus Euclidean distance for graph construction, intrinsic-dimension estimation, B = 5000 permutations on borderline statistics) . The central empirical finding is a qualitative reorganization of the geometric encoding of identity across the instruction-tuning boundary: in the base-weight Gemma-4-E4B, the identity fingerprint is encoded predominantly in the direction of hidden-state vectors (the Wasserstein separation is 0.034 with permutation p = 0 .002 under angular k-NN, where the norm is neutralized); in the multimodal instruction-tuned Gemma-4-E4B-it the fingerprint migrates into the magnitude: the separation collapses under angular k-NN (p = 0 .439) but survives under Euclidean k-NN (p = 0 .047, refined to p = 0 .042 at B = 5000), and the mean norm of the first generated state is markedly lower under the identity prompt (∥v1 ∥ = 138 .9) than under both the generic (211 .5) and vanilla (195 .3) conditions, in inversion of the relationship in the base model. This direction-to-magnitude reorganization is specific to the multimodal instruction-tuning regime: it is absent under RL distillation (separations track length, not content) and under SFT instruction-tuning (no separations) . A teacher-forced control quantifies that ∼ 30% of the free-running cosine signal is prompt-driven (vs ∼ 70% content-driven) . We position the methodological combination, W1 on edge-wise distributions of Ollivier-Ricci curvature on k-NN trajectory graphs, as a contribution of independent interest.  \n1 Introduction  \n1.1 Context  \nTransformer language models [Vaswani et al. , 2017] produce fluent text conditional on a prompt. Substantial effort has been devoted to characterizing what they output under various conditioning strategies and where specific features reside in their parameters [Hewitt and Manning, 2019 , Tenney et al. , 2019 , Belinkov and Glass, 2019 , Elhage et al. , 2021] . Comparatively less attention has been devoted to the geometry of the hidden-state trajectories these models produce during autoregressive generation: the sequence v 1 , v2 , . . . , vN ∈ H of internal representations produced at successive generation steps, viewed as a discrete path in the hidden-state space H ⊂ RD .  \nRecent work begins to close this gap. Intrinsic-dimensionality profile","cbCaihVlHv222G4P","https://ap.wps.com/l/cbCaihVlHv222G4P","pdf",749071,1,16,"English","en",105,"# Abstract\n# Introduction\n## Context\n## Research question\n## Theoretical framing (minimal)","[{\"question\":\"What is the main research question of the paper?\",\"answer\":\"The paper asks whether identity-specifying system prompts induce a statistically distinguishable geometric fingerprint in transformer hidden-state trajectories beyond prompt length and generated content, and how that fingerprint varies by post-training regime.\"},{\"question\":\"Which models and post-training regimes are compared?\",\"answer\":\"Four open-weight transformer language models are examined: Gemma-4-E4B base (no training), Gemma-4-E4B-it (multimodal RLHF), DeepSeek-R1-DistillQwen-7B (RL distillation), and Qwen2.5-7B-Instruct (supervised instruction-tuning).\"},{\"question\":\"What is the key empirical finding about where the identity information is encoded?\",\"answer\":\"The identity fingerprint reorganizes across the instruction-tuning boundary: it is primarily direction-coded in the base model but migrates to magnitude-coded encoding in the multimodal instruction-tuned model, while separations are absent under RL distillation and SFT instruction-tuning.\"}]",1784201393,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"from-direction-to-magnitude-how-multimodal-instruction-tuning-reorganizes-the-geometric-encoding-of-identity-specifying-prompts-in-transformer-hidden-states","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/from-direction-to-magnitude-how-multimodal-instruction-tuning-reorganizes-the-geometric-encoding-of-identity-specifying-prompts-in-transformer-hidden-states/85149/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main research question of the paper?","Question",{"text":75,"@type":76},"The paper asks whether identity-specifying system prompts induce a statistically distinguishable geometric fingerprint in transformer hidden-state trajectories beyond prompt length and generated content, and how that fingerprint varies by post-training regime.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which models and post-training regimes are compared?",{"text":80,"@type":76},"Four open-weight transformer language models are examined: Gemma-4-E4B base (no training), Gemma-4-E4B-it (multimodal RLHF), DeepSeek-R1-DistillQwen-7B (RL distillation), and Qwen2.5-7B-Instruct (supervised instruction-tuning).",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key empirical finding about where the identity information is encoded?",{"text":84,"@type":76},"The identity fingerprint reorganizes across the instruction-tuning boundary: it is primarily direction-coded in the base model but migrates to magnitude-coded encoding in the multimodal instruction-tuned model, while separations are absent under RL distillation and SFT instruction-tuning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":28,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]