[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86355-en":3,"doc-seo-86355-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86355,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","FunOCLUST Clustering Functional Data with Outliers","Functional data clustering faces challenges from infinite-dimensional structure and heightened sensitivity to atypical observations. The document proposes funOCLUST, an extension of the OCLUST algorithm tailored to functional settings. By combining model-based clustering with an iterative trimming mechanism, the method clusters curves while removing outlying trajectories. Evaluation on simulated and real-world functional datasets shows strong performance for both clustering quality and outlier identification.","arXiv :2508 .00110v2 [ stat .ML] 13 Jul 2026  \nfunOCLUST: Clustering Functional Data with Outliers  \nKatharine M. Clark 1 and Paul D. McNicholas 2  \n1  \nDepartment of Mathematics & Statistics, Trent University, Ontario, Canada.  \n2 Department of Mathematics & Statistics, McMaster University, Ontario, Canada.  \nAbstract  \nFunctional data present unique challenges for clustering due to their infinite-dimensional nature and potential sensitivity to outliers. An extension of the OCLUST algorithm to the functional setting is proposed to address these issues. The approach leverages the OCLUST framework, creating a robust method to cluster curves and trim outliers. The methodology is evaluated on both simulated and real-world functional datasets, demonstrating strong performance in clustering and outlier identification.  \nKeywords: OCLUST, curves, mixture models, outliers.  \n1 Introduction  \nFunctional data analysis presents a unique challenge due to the nature of the data. While functions are observed at discrete time points, they exist in an infinite-dimensional function space, which requires methods that respect their smoothness, continuity, and underlying structure rather than treating them as simple multivariate observations. A natural secondary aim of functional data analysis is to group similar curves or trajectories using cluster analysis, e.g. clustering girls’ versus boys’ growth curves and weather stations based on recorded temperatures (Ramsay and Silverman, 2005) .  \nClustering is a subtype of classification, i.e., unsupervised classification, which aims to group similar sets of observations together without any prior knowledge of the cluster memberships. While non-parametric methods exist, e.g. , hierarchical and k-means clustering, model-based clustering employs a parametric approach. In general, the density of a finite mixture model is  \nG  \nf (x | ϑ) =X πgfg (x | θg) , (1)  \ng=1  \nwhere ϑ = {π1 , . . . , πG , θ 1 , . . . θG } , πg > 0 is the gth mixing proportion with P πg = 1 , and fg (x | θg) is the gth component density with parameters θg . Typically, each component corresponds to a cluster (see McNicholas, 2016, for a discussion) . While each cluster can be modelled with various component distributions, the Gaussian distribution remains popular  \ndue to its simplicity. The parameters and cluster memberships are estimated by maximizing the log-likelihood, often with the expectation-maximization (EM) algorithm (Dempster et al. , 1977) .  \nWhile model-based clustering methods for functional data have been developed (e.g. , James and Sugar, 2003 ; Bouveyron and Jacques, 2011), they remain sensitive to noisy observations. Alternatives with more robust distributions (Anton and Smith, 2023 ; AmovinAssagba et al. , 2022) and trimming approaches (Rivera-García et al. , 2019) have been proposed, but these methods remain focused on subspace-specific clustering. While subspace methods can be highly effective, there are settings in which the clustering structure requires the entire functional domain, making dimension reduction less desirable.  \nNon-parametric trimming methods do exist (e.g. , Garcia-Escudero and Gordaliza, 2005), but there is a scarcity of trimming algorithms within the model-based functional data clustering domain. The proposed algorithm, funOCLUST, combines trimming with model-based clustering to detect outliers and cluster similar functions simultaneously while retaining information from the full functional representation.  \n2 Preliminaries  \n2.1 OCLUST Algorithm  \nIn the multivariate normal data domain, Clark and McNicholas (2024) develop the OCLUST algorithm, which simultaneously clusters data and trims outliers. Data are modelled with mixtures of Gaussian distributions, and outliers are trimmed iteratively one-by-one. A subset log-likelihood is defined to be the log-likelihood of the data with a single data point removed, and there are n such subsets. After each outlier removal, the distribution of the ","cbCaihqNN8XJ3zea","https://ap.wps.com/l/cbCaihqNN8XJ3zea","pdf",2761731,4,1,20,"English","en",105,"# Introduction\n## Functional data clustering and robustness\n## Model-based clustering and limitations\n# Preliminaries\n## OCLUST algorithm\n## Functional decomposition","[{\"question\":\"What problem does funOCLUST address in functional data clustering?\",\"answer\":\"It targets clustering difficulties caused by the infinite-dimensional nature of functional data and the sensitivity of clustering to outliers.\"},{\"question\":\"How does funOCLUST extend OCLUST for functional data?\",\"answer\":\"It leverages the OCLUST framework to build a robust clustering approach that trims outliers while clustering curves using a functional-data adaptation.\"},{\"question\":\"How are the method’s performance claims evaluated?\",\"answer\":\"The approach is assessed on both simulated datasets and real-world functional datasets, focusing on clustering effectiveness and outlier identification accuracy.\"}]",1784210800,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"funoclust-clustering-functional-data-with-outliers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/funoclust-clustering-functional-data-with-outliers/86355/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does funOCLUST address in functional data clustering?","Question",{"text":75,"@type":76},"It targets clustering difficulties caused by the infinite-dimensional nature of functional data and the sensitivity of clustering to outliers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does funOCLUST extend OCLUST for functional data?",{"text":80,"@type":76},"It leverages the OCLUST framework to build a robust clustering approach that trims outliers while clustering curves using a functional-data adaptation.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the method’s performance claims evaluated?",{"text":84,"@type":76},"The approach is assessed on both simulated datasets and real-world functional datasets, focusing on clustering effectiveness and outlier identification accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]