[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81545-en":3,"doc-seo-81545-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81545,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Cluster and then Embed: A Modular Approach for Visualization","Dimensionality reduction methods like t-SNE and UMAP are widely used to visualize data that may contain latent clustered structure. They often produce embeddings where clusters appear well separated and local neighborhoods are preserved, yet they can significantly distort the global geometry. A modular alternative is proposed: first cluster the data, then embed each cluster independently, and finally align the cluster embeddings to form a coherent global view. Experiments on synthetic and real datasets show competitive performance with improved transparency.","arXiv :2509 .03373v2 [ cs .LG] 10 Jul 2026  \nCluster and then Embed: A Modular Approach for Visualization  \nElizabeth Coda 1 , Ery Arias-Castro 1,2 , and Gal Mishne2  \n1 Department of Mathematics, University of California, San Diego  \n2 Halıcıo˘glu Data Science Institute, University of California, San Diego  \nAbstract  \nDimensionality reduction methods such as t-SNE and UMAP are popular methods for visualizing data with a potential (latent) clustered structure. They are known to group data points at the same time as they embed them, resulting in visualizations with well-separated clusters that preserve local information well. However, t-SNE and UMAP also tend to distort the global geometry of the underlying data. We propose a more transparent modular approach that first clusters the data, then embeds each cluster, and finally aligns the clusters to obtain a global embedding. We demonstrate this approach on several synthetic and real-world datasets and show that it is competitive with existing methods, while being much more transparent.  \n1 Introduction  \nVisualization is one of the most commonly used tools in exploratory data analysis (EDA) [9] . However, when the data is high-dimensional, direct visualization of the data is not possible and dimensionality reduction techniques must be used to obtain a low-dimensional visualization.  \nIn this paper, we assume the data are of the form {xi} ⊂ Rd. We seek an embedding {yi} ⊂ Rm with m ≤ d so that if a clustered structure is present in the data, the embedding uncovers the structure and provides visualization of that structure. Typically, m = 2 and we will assume this is the case for the remainder of the paper, though results can easily be extended tom > 2. As stated, the problem is ill-defined as it is not clear what exactly is to be optimized.  \nThe visualization of clustered data is certainly not a new problem. Early work includes that of Shepard [36], who proposed the practice of imposing the clusters obtained from a hierarchical clustering algorithm onto an embedding. Shepard explains that imposing hierarchical clustering onto the spatial representation provides more information than a hierarchical cluster tree alone (e.g., that two clusters are closer to one another than a third cluster) . While his proposal is based on small datasets (n = 16 in his motivating example), the practice of imposing clustering on an embedding is very common. The visualization of single-cell RNA sequencing data in Bioinformatics is a case in point, as data of that sort are typically expected to be clustered [5] . Such data are often visualized by first producing an embedding using, e.g., t-SNE [45] or UMAP [28], and the points are then colored according to known labels or the labeling of some clustering method, e.g., Louvain [4], Leiden [43], or DBSCAN [11] . Similarly, in machine learning in the task of classification, it is common to visualize high-dimensional data such as images or the hidden layers of a neural network with the embedded points colored according to the class label.  \nMany classical approaches for constructing an embedding seek to preserve some measure of dissimilarity between points, often the Euclidean distance. Examples include classical scaling (CS), also sometimes referred to as multidimensional scaling (MDS), which coincides with principal component analysis (PCA) when the dissimilarity measure is the Euclidean distance; and Isomap [42],  \n2  \nwhich approximates the geodesic distances between points under the assumption that they lie on a smooth submanifold, and then embeds the points using CS with these distances. However, these methods for dimension reduction were not explicitly designed for the visualization of clustered data, and tend not to separate the clusters well. In particular, when embedding a configuration of points in Rd into R2 or R3 there may not be enough space in the low-dimensional embedding space to accommodate all points at a given distance scale. As a res","cbCaikTNlqv0oE8K","https://ap.wps.com/l/cbCaikTNlqv0oE8K","pdf",24814662,2,1,28,"English","en",105,"# Introduction\n## Visualization for High-Dimensional Data\n## Classical and Nonlinear Dimensionality Reduction\n## Neighborhood-Preserving Methods for Cluster Separation","[{\"question\":\"Why do t-SNE and UMAP sometimes fail to preserve the global geometry of data?\",\"answer\":\"They excel at separating clusters and preserving local neighborhoods, but they can distort the global geometry of the underlying data, leading to misleading global relationships in the low-dimensional view.\"},{\"question\":\"What is the core idea of the proposed modular approach?\",\"answer\":\"The method clusters the data first, then embeds each cluster separately, and finally aligns the embedded clusters to obtain a global embedding that is more interpretable.\"},{\"question\":\"How is cluster visualization related to early exaggeration in t-SNE?\",\"answer\":\"The document notes that changing the t-SNE early exaggeration parameter yields a range of embeddings: small values emphasize discrete cluster structure and higher kNN recall, while larger values increase attractive forces and highlight connections between clusters.\"}]",1784174203,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cluster-and-then-embed-a-modular-approach-for-visualization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cluster-and-then-embed-a-modular-approach-for-visualization/81545/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do t-SNE and UMAP sometimes fail to preserve the global geometry of data?","Question",{"text":75,"@type":76},"They excel at separating clusters and preserving local neighborhoods, but they can distort the global geometry of the underlying data, leading to misleading global relationships in the low-dimensional view.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea of the proposed modular approach?",{"text":80,"@type":76},"The method clusters the data first, then embeds each cluster separately, and finally aligns the embedded clusters to obtain a global embedding that is more interpretable.",{"name":82,"@type":73,"acceptedAnswer":83},"How is cluster visualization related to early exaggeration in t-SNE?",{"text":84,"@type":76},"The document notes that changing the t-SNE early exaggeration parameter yields a range of embeddings: small values emphasize discrete cluster structure and higher kNN recall, while larger values increase attractive forces and highlight connections between clusters.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]