[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81985-en":3,"doc-seo-81985-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81985,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Riemannian Geometry for Pre-trained Language Model Embeddings","Understanding the geometric structure of pretrained language model embeddings matters for interpretability and safety. This work tests whether sentence-level classification signal resides in the Riemannian geometry of contextual token embeddings by extracting per-token pullback metrics from an analytical Jacobian and aggregating them with a Fréchet mean on the symmetric positive definite (SPD) manifold, termed Riemannian Mean Pooling (RMP). On CoLA, CREAK, and RTE, RMP improves over Euclidean mean pooling, while on FEVER-Symmetric it stays at chance. Ablations attribute the gain to geometric aggregation rather than learned manifold structure, with extra signal on CREAK.","Riemannian Geometry for Pre-trained Language Model Embeddings  \nSzczepan Konior 1 Alexandre Quemy2 Przemysław Klocek 1  \nBartłomiej Sobieski3 ,4 Grégoire Cattan 1  \n1IBM Automation and AI, Krakow, Poland 2Hother, Krakow, Poland  \n3University of Warsaw 4 Centre for Credible AI, Warsaw University of Technology {[name.surname}@ibm.com](name.surname}@ibm.com)  \n[alexandre@hother.io](alexandre@hother.io)  \n[sobieski.bartlomiej.jan@gmail.com](sobieski.bartlomiej.jan@gmail.com)  \narXiv :2607 .07047v2 [ cs .CL] 10 Jul 2026  \nAbstract  \nUnderstanding the geometric structure of pretrained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting pertoken pullback metrics from a learned encoder’s analytical Jacobian and aggregating them with the Fréchet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP) . Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.  \n1 Introduction  \nPre-trained language models achieve remarkable success in natural language understanding, yet the geometric structure of their internal representations remains an open challenge. Most analyses treat token embeddings as points in flat Euclidean space, but three lines of evidence motivate departing from that assumption: hierarchical structure in language admits lower-distortion embeddings in negatively curved spaces (Nickel and Kiela, 2017, 2018 ; Sala et al., 2018 ; Park et al., 2025) and is reflected in non-standard inner products on the embedding space (Park et al., 2024 ; Marks and Tegmark, 2024); contextualized  \nrepresentations in BERT, ELMo, and GPT-2 are strongly anisotropic, occupying narrow cones not invariant under Euclidean transformation (Ethayarajh, 2019 ; Kataiwa et al., 2025); and token embedding spaces may fail to form smooth manifolds globally, even if local-manifold-like structure still supports geometric analysis at the neighbourhood level (Robinson et al., 2025) .  \nThis work investigates whether sentence-level classification signal lives in the Riemannian geometry of token embeddings. We address three questions: (1) Do embeddings exhibit local geometric structure accessible via pullback metrics?  \n(2) Does this geometric structure carry signal beyond what flat Euclidean aggregation extracts?  \n(3) Which component of a geometric aggregation pipeline is responsible for any observed gain: the encoder, the metric, or the aggregation?  \nWe evaluate our approach, Riemannian Mean Pooling (RMP), on four diverse NLP tasks: fact verification (FEVER-Symmetric), textual entailment (RTE), grammatical acceptability (CoLA), and commonsense reasoning (CREAK) . FEVERSymmetric is included specifically as a negative control: it was constructed to remove lexical and annotation artifacts (Schuster et al., 2019), and we expect any method relying solely on claim-surface geometry to perform at chance.  \nAny differentiable encoder yields a pullback metric through its Jacobian, but the metric’s geometric interpretability depends on the training objective. We use Intrinsic Green’s Learning (IGL) (Quemy, 2026)1 : a closed-form kernel readout replaces the learned decoder, constraining the encoder to coordinates aligned with the smooth structure of the data manifold rather t","cbCaioqG1zPqMcEn","https://ap.wps.com/l/cbCaioqG1zPqMcEn","pdf",892693,6,1,15,"English","en",105,"# Introduction\n# Related work\n# Contributions","[{\"question\":\"What is Riemannian Mean Pooling (RMP) in this work?\",\"answer\":\"RMP extracts per-token pullback metrics from an encoder’s analytical Jacobian, aggregates them with a Fréchet mean on the SPD manifold, and performs classification in the tangent space at the population Riemannian mean.\"},{\"question\":\"How does RMP perform compared with Euclidean mean pooling?\",\"answer\":\"RMP outperforms Euclidean mean pooling on three signal-bearing datasets (CoLA, CREAK, RTE). On FEVER-Symmetric, designed to remove lexical and annotation artifacts, the method stays at chance.\"},{\"question\":\"What do the ablations reveal about where the performance gain comes from?\",\"answer\":\"A randomly initialised encoder with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, indicating the gain mainly comes from geometric aggregation. The trained encoder adds extra signal specifically on CREAK.\"}]",1784177419,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"riemannian-geometry-for-pre-trained-language-model-embeddings","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/riemannian-geometry-for-pre-trained-language-model-embeddings/81985/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Riemannian Mean Pooling (RMP) in this work?","Question",{"text":76,"@type":77},"RMP extracts per-token pullback metrics from an encoder’s analytical Jacobian, aggregates them with a Fréchet mean on the SPD manifold, and performs classification in the tangent space at the population Riemannian mean.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does RMP perform compared with Euclidean mean pooling?",{"text":81,"@type":77},"RMP outperforms Euclidean mean pooling on three signal-bearing datasets (CoLA, CREAK, RTE). On FEVER-Symmetric, designed to remove lexical and annotation artifacts, the method stays at chance.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the ablations reveal about where the performance gain comes from?",{"text":85,"@type":77},"A randomly initialised encoder with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, indicating the gain mainly comes from geometric aggregation. The trained encoder adds extra signal specifically on CREAK.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]