[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85071-en":3,"doc-seo-85071-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85071,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","TypeProbe Recovering Type Representations from Hidden States of Pre-trained Code Models","State-of-the-art code models perform well yet their internal encoding of type information is not well understood. TypeProbe investigates residual streams of pretrained code models using a parallel dataset of Java and Python programs. It finds cross-lingual type representations can arise even from untyped code. Linear probing is used to test whether hidden states encode result types implied by typed function application across languages, and the discovered structure shows partial robustness to lexical perturbations and cross-language syntactic variation.","arXiv :2607 .08339v 1 [ cs .CL] 9 Jul 2026  \nTypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models  \nGiuliano Gorgone 1[0009−0004−0827−592X], Fausto Carcassi 1[0000−0001−6843−1737]  \nILLC, University of Amsterdam, Amsterdam, The Netherlands [giuliano.gorgone@student.uva.nl](giuliano.gorgone@student.uva.nl) , [f.carcassi@uva.nl](f.carcassi@uva.nl)  \nAbstract. State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly understood. We probe the residual streams of pretrained code models for internal type representations using a parallel dataset of Java and Python code examples. Our results show that cross-lingual type representations emerge even from untyped code. Moreover, we test whether hidden states linearly encode the result type implied by typed function application by training probes on one language to infer argument and result types in the other. Finally, we find that this structure is partly robust to lexical perturbations and cross-language syntactic variations.  \nTo the best of our knowledge, prior work on interpretability of code models has not directly targeted formal type semantics or cross-lingual type representations. We release our code and datasets 1 .  \nKeywords: Code LLMs · Linear Probing · Type Semantics · Cross-lingual Representations · Adversarial Renaming · Lexical Interference  \n1 Introduction  \nState-of-the-art code models excel at next-token prediction but can still violate formal type-system constraints. For instance, in generated TypeScript, nearly 94% of compilation errors stem from type-check failures [10] . Type-constrained decoding [10] mitigates this by enforcing language-specific rules externally, but such approaches assume that models cannot internalize formal type information and require ad hoc machinery for each language.  \nDiagnostic probing has become a standard tool for studying abstract information encoded in neural representations [1,7 ,4], and prior work on pretrained code models shows that they capture syntax, identifiers, and namespaces more readily than complex semantic properties [13] . At the same time, recent activationsteering results suggest that code models contain a controllable type-prediction mechanism shared across languages [9], motivating a more direct investigation of how such information is represented internally. In natural language, multilingual representations often exhibit approximately aligned manifolds across languages  \n1 [https://github.com/anticleiades/TypeProbe](https://github.com/anticleiades/TypeProbe)  \n2 G. Gorgone, F. Carcassi  \n[3,8]; here we ask whether an analogous structure emerges for type information in code models.  \nIn this work, we use layer-wise linear probes to test whether pretrained code models encode type information in their residual streams, where this information is localized across depth, and whether it transfers across Java and Python.  \n2 Interpreting Code Models  \nWe investigate the internal logic of code models by addressing three primary research questions. First, we examine representation: do pretrained code models develop linearly decodable representations of type-level semantics, and in which transformer layers does this information emerge? Second, we explore invariance and robustness: is this representation invariant to language syntax and robust against lexical-level adversarial perturbations? More specifically, do models prioritize formal type constraints over superficial heuristics when identifier names explicitly contradict their true underlying types? Finally, we assess inference: beyond type prediction, can models leverage these internal representations to resolve data flow and perform typed function application?  \n2.1 Experimental Setup  \nDataset and Design. We construct a programmatically generated dataset with three 90K-example partitions: Java (Java), pyUnt (Python without type annotations), and pyTag (Pyt","cbCaihUt34vQr9MP","https://ap.wps.com/l/cbCaihUt34vQr9MP","pdf",901303,4,1,18,"English","en",105,"# Introduction\n# Interpreting Code Models\n## Experimental Setup\n### Dataset and Design\n### Type Systems","[{\"question\":\"What is the main goal of TypeProbe?\",\"answer\":\"TypeProbe aims to recover and analyze internal type representations in pretrained code models by probing their residual streams, with a focus on cross-lingual behavior between Java and Python.\"},{\"question\":\"How is the dataset constructed to avoid lexical shortcuts?\",\"answer\":\"The dataset includes Java and Python variants with and without type annotations, while identifiers and literals are independently randomized per partition so that operationally equivalent examples do not share exact names or values across languages.\"},{\"question\":\"What transfer and robustness effects does the work report?\",\"answer\":\"The work reports that cross-lingual type representations emerge even from untyped code. Probes trained in one language can infer argument and result types in the other, and the learned structure is partly robust to lexical perturbations and syntactic variations across languages.\"}]",1784200818,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"typeprobe-recovering-type-representations-from-hidden-states-of-pre-trained-code-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/typeprobe-recovering-type-representations-from-hidden-states-of-pre-trained-code-models/85071/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of TypeProbe?","Question",{"text":75,"@type":76},"TypeProbe aims to recover and analyze internal type representations in pretrained code models by probing their residual streams, with a focus on cross-lingual behavior between Java and Python.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the dataset constructed to avoid lexical shortcuts?",{"text":80,"@type":76},"The dataset includes Java and Python variants with and without type annotations, while identifiers and literals are independently randomized per partition so that operationally equivalent examples do not share exact names or values across languages.",{"name":82,"@type":73,"acceptedAnswer":83},"What transfer and robustness effects does the work report?",{"text":84,"@type":76},"The work reports that cross-lingual type representations emerge even from untyped code. Probes trained in one language can infer argument and result types in the other, and the learned structure is partly robust to lexical perturbations and syntactic variations across languages.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]