[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128234-en":3,"doc-seo-128234-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128234,2336475104736,"Quinn","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222",8,"Research & Report","Contributions on dimensionality reduction and interpretable machine learning - PhD thesis","This PhD thesis is organized into two parts: dimensionality reduction for large datasets and interpretable machine learning. The first part introduces algorithms implementing multidimensional scaling (MDS) for settings where classical MDS is infeasible due to excessive memory and time costs. Six non-standard partition-based methods are compared through simulation, and an open-source R package (bigmds) is provided. The second part develops interpretability approaches, including SurvLIMEpy for survival analysis and a Shapley-value method for functional data regression.","Contributions on dimensionality reduction and interpretable machine learning  \nCristian Pachón-García  \nAdvisor: Pedro Delicado  \nThesis presented for the degree of Doctor  \nPhD programme in Statistics and Operations Research  \nDepartment of Statistics and Operations Research  \nBarcelona, December 2024  \nA la Maula, la Natàlia i la Bibi.  \nAcknowledgements  \nEn primer lugar, me gustaría mostrar mi más sincero y profundo agradecimiento a Pedro, mi director, mentor y guía durante este viaje. Muchas gracias por todo lo que has hecho por mí, tanto a nivel personal como profesional. Eres una fuente inagotable de conocimiento y has sabido transmitirmeparte de él, he aprendido mucho trabajando contigo. Uno de los momentos más difíciles para mí fue cuando recibimos las segundas correcciones del paper de MDS. A pesar que era parte de mi doctorado, te pusiste a trabajar a mi lado para que el paper tuviera un buen fin. Sin ti, este proceso habría sido mucho más difícil de llevar a cabo.  \nEn segon lloc, gràcies per tot, Sòcia. Tot i les conseqüències que tenia deixar una feina estable peruna beca de doctorat, em vas empènyer a fer-ho. Han estat 4 anys plens de canvis a la nostra vida, quins 4 anys! Un doctorat és una bona barreja d’emocions, gràcies per haver-les escoltat i haver-me permès compartir-les amb tu. Haver acabat aquest document amb la teva ajuda ha estat molt bonic. Has invertit moltes hores i esforços en editar aquesta tesi, que no hauria quedat tan bé sense tu, així que moltes gràcies, amor meu.  \nTambién me gustaría dar las gracias a Verónica. Aceptaste el liderazgo del proyecto de SurvLIMEpy a pesar de todo el trabajo que tenías. Siempre tuviste tiempo para mí . Me ha gustado mucho conocerte y haber compartido contigo este proyecto.  \nPor otro lado, quiero dar las gracias a Carlos. Compañero, todo un placer haber compartido contigo el proyecto de SurvLIMEpy. Eres un gran profesional, muy comprometido con el trabajo. Siento que el doctorado me ha llevado a hacer un buen amigo. Admiro la capacidad de autoaprendizaje que tienes.  \nFinally, and more formally, I would like to express my gratitude to the Department of Statistics and Operational Research at UPC. I would also like to thank the funding institutions and projects: AGAUR, La Agencia Estatal de Investigación and UPC.  \nAbstract  \nThis thesis is divided into two parts. The first one is devoted to dimensionality reduction for large datasets, while the second one focuses on the field of interpretable machine learning. Part of the material presented in this thesis has been published either in journals or workshops. The work of Delicado and Pachón-García (2024b) is used to elaborate the first part of this thesis, which are Chapters 1, 2 and 3. Regarding Chapter 5, the original publication is Pachón-García et al. (2024) and the material of Section 5.4 is based on Hernández-Pérez et al. (2024) . Finally, Chapter 6 is based on Delicado and Pachón-García (2024a), which is a preprint version to be submitted to a journal.  \nTo begin with, we present a set of algorithms implementing multidimensional scaling (MDS) for large data sets. MDS is a family of dimensionality reduction techniques using a n × n distance matrix as input, where n is the number of individuals, and producing a low dimensional configuration: a n × r matrix with r ≪ n. When n is large, MDS is unaffordable with classical MDS algorithms because their extremely large memory and time requirements.  \nWe compare six non-standard algorithms intended to overcome these difficulties. They are based on the central idea of partitioning the data set into small pieces, where classical MDS methods can work. Two of these algorithms are original proposals. In order to check the performance of the algorithms as well as to compare them, we have done a simulation study. In addition, an open-source R package implementing the algorithms has been created, bigmds.  \nRegarding the field of machine learning (ML), it is worth noting that ","cbCaiogbCcqLFBOm","https://ap.wps.com/l/cbCaiogbCcqLFBOm","pdf",8957866,2,1,108,"English","en",105,"# Abstract\n## Dimensionality reduction for large datasets\n## Interpretable machine learning and survival analysis\n## SurvLIMEpy and model explainability\n## Functional data interpretability using Shapley values","[{\"question\":\"How is the thesis structured and what are its two main research areas?\",\"answer\":\"The thesis is divided into two parts: dimensionality reduction for large datasets and interpretable machine learning. The focus shifts from MDS algorithms to interpretability methods for ML models.\"},{\"question\":\"What challenge motivates the dimensionality reduction part of the work?\",\"answer\":\"Classical MDS becomes unaffordable for large datasets because it requires extremely large memory and time when using the n×n distance matrix. The thesis targets large-data scalability.\"},{\"question\":\"What does SurvLIMEpy contribute to interpretable machine learning?\",\"answer\":\"SurvLIMEpy is an open-source Python package that implements SurvLIME to compute local feature importance for survival analysis models. It includes visualization tools and is evaluated via simulation and survival-model comparisons using the SEER database.\"}]","Contributions on dimensionality reduction and interpretable machine learning - PhD thesis | PDF",1785945976,272,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"contributions-on-dimensionality-reduction-and-interpretable-machine-learning-phd-thesis","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/contributions-on-dimensionality-reduction-and-interpretable-machine-learning-phd-thesis/128234/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How is the thesis structured and what are its two main research areas?","Question",{"text":76,"@type":77},"The thesis is divided into two parts: dimensionality reduction for large datasets and interpretable machine learning. The focus shifts from MDS algorithms to interpretability methods for ML models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What challenge motivates the dimensionality reduction part of the work?",{"text":81,"@type":77},"Classical MDS becomes unaffordable for large datasets because it requires extremely large memory and time when using the n×n distance matrix. The thesis targets large-data scalability.",{"name":83,"@type":74,"acceptedAnswer":84},"What does SurvLIMEpy contribute to interpretable machine learning?",{"text":85,"@type":77},"SurvLIMEpy is an open-source Python package that implements SurvLIME to compute local feature importance for survival analysis models. It includes visualization tools and is evaluated via simulation and survival-model comparisons using the SEER database.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]