[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82954-en":3,"doc-seo-82954-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82954,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes","Categorical data clustering depends heavily on distance metrics that quantify dissimilarity between objects, yet many methods handle nominal and ordinal attributes identically and ignore ordinal order information. This work models the intrinsic differences and connections between nominal and ordinal values graph-like and introduces a unified intra-attribute distance metric that preserves ordinal ordering. A new learning clustering algorithm jointly learns intra-attribute distance weights and data partitions in one paradigm, avoiding suboptimal solutions and improving performance over existing approaches.","Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes  \nYiqun Zhang, Member, IEEE and Yiu-ming Cheung, Fellow, IEEE  \narXiv :2607 .05464v 1 [ cs .LG] 6 Jul 2026  \nAbstract—The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two objects. However, most of the existing clustering methods treat the two categorical subtypes, i.e. nominal and ordinal attributes, in the same way when calculating the dissimilarity without considering the relative order information of the ordinal values. Moreover, there would exist interdependence among the nominal and ordinal attributes, which is worth exploring for indicating the dissimilarity. This paper will therefore study the intrinsic difference and connection of nominal and ordinal attribute values from a perspective akin to the graph. Accordingly, we propose a novel distance metric to measure the intra-attribute distances of nominal and ordinal attributes in a unified way, meanwhile preserving the order relationship among ordinal values. Subsequently, we propose a new clustering algorithm to make the learning of intra-attribute distance weights and partitions of data objects into a single learning paradigm rather than two separate steps, whereby circumventing a suboptimal solution. Experiments show the efficacy of the proposed algorithm in comparison with the existing counterparts.  \nIndex Terms—Categorical data clustering, nominal-and-ordinal attribute, intra-attribute distance, learnable weighting.  \n~~ ~~ ✦ ~~ ~~  \n1 INTRODUCTION  \nWIDESPREAD categorical data can be easily collected  \nfrom questionnaires, medical scales, scoring systems, and so on [1] . As one of the most widely used machine learning and pattern recognition techniques, clustering that partitions data objects into homogeneous groups in unsupervised environment [2], [3] has been commonly adopted for the analysis of categorical data [4], [5] . In order to better discover homogeneous clusters, weighting attributes according to their importance to the clustering task [6] is adopted by many existing clustering algorithms [7], [8],[9], [10], [11] . Since weighting an attribute is equivalent to uniformly weighting all the intra-attribute distances measured on this attribute, these algorithms are actually based on the hypothesis that all the intra-attribute distances are well defined, which is reasonable for numerical data with well-defined distance measure [12] . However, for categorical data whose distance measure is generally not well-defined, uniformly weighting the intra-attribute distances is surely unreasonable [13] . To solve this problem, most existing methods focus on exploring appropriate distance measures [14],[15] and attribute weighting mechanisms [11] .  \nSuccessful attempts in exploring appropriate distance measures include Lin’s [16] similarity measure, coupled [17] similarity metric, association-based [18], Ahmad’s [19], context-based [20], [21], and Jia’s [22] distance metrics. The above-mentioned measures define intra-attribute distances according to the possible value statistics, e.g., the occurrence frequencies and conditional occurrence probabilities. Lin’s measure computes the cumulative entropy of a range of ordered possible values (i.e., the adjacent possible values {good, neutral, bad} of an ordinal attribute with possible values {very-good, good, neutral, bad, very-bad}) to indicate the corresponding intra-attribute distance (i.e., the distance between good and bad) with preserving the order  \nTABLE 1: Fragment of Lymphography data set.  \n\n| No. | Attribute 1\u003Cbr>(enlarge) | Attribute 2\u003Cbr>(form) | Class\u003Cbr>(diagnosis) |\n| --- | --- | --- | --- |\n| 1 ↑ |  | non-special normal |  |\n| 2 | ↑ | vesicles | fibrosis |\n| 3 | ↑↑ | vesicles | fibrosis |\n| 4 | ↑↑↑ | chalices | metastases |\n| 5 | ↑↑↑ | chalices | malign |\n| 6 | ↑↑↑↑ | vesicles | malign |\n\nrelationship, whic","cbCaiknH5BF1t2a7","https://ap.wps.com/l/cbCaiknH5BF1t2a7","pdf",622050,2,1,16,"English","en",105,"# Introduction\n## Distance metrics for categorical clustering\n## Challenges with mixed nominal and ordinal attributes\n## Proposed approach overview","[{\"question\":\"Why do existing categorical clustering methods often underperform for ordinal attributes?\",\"answer\":\"They usually compute dissimilarity treating nominal and ordinal attribute subtypes the same way, thereby ignoring the relative order information carried by ordinal values.\"},{\"question\":\"How does the proposed method handle nominal and ordinal attributes differently in measuring dissimilarity?\",\"answer\":\"It studies both the intrinsic difference and the connection between nominal and ordinal values, then proposes a unified distance metric that preserves the order relationship among ordinal values while still addressing nominal/ordinal interdependence.\"},{\"question\":\"What is the main improvement in the new clustering algorithm compared with prior approaches?\",\"answer\":\"It jointly learns intra-attribute distance weights and cluster partitions in a single learning paradigm rather than using two separate steps, which is intended to avoid suboptimal solutions; experiments show better efficacy than existing methods.\"}]",1784184307,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learnable-weighting-of-intra-attribute-distances-for-categorical-data-clustering-with-nominal-and-ordinal-attributes","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learnable-weighting-of-intra-attribute-distances-for-categorical-data-clustering-with-nominal-and-ordinal-attributes/82954/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing categorical clustering methods often underperform for ordinal attributes?","Question",{"text":75,"@type":76},"They usually compute dissimilarity treating nominal and ordinal attribute subtypes the same way, thereby ignoring the relative order information carried by ordinal values.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method handle nominal and ordinal attributes differently in measuring dissimilarity?",{"text":80,"@type":76},"It studies both the intrinsic difference and the connection between nominal and ordinal values, then proposes a unified distance metric that preserves the order relationship among ordinal values while still addressing nominal/ordinal interdependence.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main improvement in the new clustering algorithm compared with prior approaches?",{"text":84,"@type":76},"It jointly learns intra-attribute distance weights and cluster partitions in a single learning paradigm rather than using two separate steps, which is intended to avoid suboptimal solutions; experiments show better efficacy than existing methods.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]