[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126279-en":3,"doc-seo-126279-105":31,"detail-sidebar-cat-0-en-105":85},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126279,2336475104736,"Quinn","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222",8,"Research & Report","Dimensionality Reduction for Machine Learning-based Argument Mining","Recent argument mining approaches train machine learning models on annotated text using high-dimensional linguistic feature vectors. This paper studies whether reducing the dimensionality of the input space can improve both efficiency and extraction quality. Experiments with SVD, PCA, and LDA on a Spanish argumentative corpus from the e-participation domain evaluate three core tasks: argumentative fragment detection, argument component classification, and argumentative relation recognition.","Dimensionality Reduction for Machine Learning-based Argument Mining  \nAndrés Segura-Tinoco and Iván Cantador  \nUniversidad Autónona de Madrid, Madrid, Spain andres.segurat@uam.es, [ivan.cantador@uam.es](ivan.cantador@uam.es)  \nAbstract  \nRecent approaches to argument mining have fo  \ncused on training machine learning algorithms from annotated text corpora, utilizing as input high-dimensional linguistic feature vectors. Differently to previous work, in this paper, we preliminarily investigate the potential benefits of reducing the dimensionality of the input data. Through an empirical study, testing SVD, PCA and LDA techniques on a new argumentative corpus in Spanish for an underexplored domain (e-participation), and using a novel, rich argument model, we show positive results in terms of both computation efficiency and argumentative information extraction effectiveness, for the three major argument mining tasks: argumentative fragment detection, argument component classification, and argumentative relation recognition. On a space with dimension around 3-4% of the number of input features, the argument mining methods are able to reach 95-97% of the performance achieved by using the entire corpus, and even surpass it in some cases.  \n1 Introduction  \nSince its origins in the late 2000s, the argument mining (AM) field has witnessed significant advances on the problem of automatically extracting structured argumentative information from text corpora (Lytos et al., 2019 ; Lawrence and Reed, 2020), which commonly entails three tasks: the identification of argumentative fragments in an input text, the split or classification of such fragments into argument components (e.g., claims and premises), and the recognition of relations (e.g., support and attack) between pairs of argument components.  \nIn particular, previous research has led to the development of effective approaches based on machine learning (ML) (Lippi and Torroni, 2015, 2016) with results almost equal to those obtained with more complex approaches, such as those based on deep learning. Hence, argumentative fragment detection (Mochales Palau and Moens, 2009a,  \n2011 ; Poudyal et al., 2016), argument component classification (Habernal and Gurevych, 2017 ; Duet al., 2017), and argument relation recognition (Duet al., 2017) have been modeled as sequence labeling problems, where, in general, each sentence 1 is represented as a vector of real-valued linguistic features and has associated certain label or class, e.g., argumentative vs. non-argumentative, and claim vs. premise. ML algorithms are thus trained with sets of labeled sentence vectors in order to predict the class of new sentences.  \nIn this context, a variety of features have been considered –ranging from lexical and morphological, to structural and syntactic, and semantic and discourse features (Stab and Gurevych, 2014 ; Aker et al., 2017 ; Habernal and Gurevych, 2017)– and, in general, approaches have dealt with feature vectors of high dimensionality.  \nTo the best of our knowledge, only a few research attempts have been made to use a subset of features (Poudyal et al., 2016 ; Du et al., 2017) . Motivated by this fact and the increasing need for more efficient (i.e., less resource-consuming) AM model building, in this paper, instead of exploring new argument-related classification algorithms, we investigate the potential benefits of reducing the dimensionality of the input data space.  \nAs an innovative research in the AM field, we report experiments conducted with the well known SVD (Beltrami, 1973 ; Stewart, 1993), PCA (Hotelling, 1933) and LDA (Fisher, 1936) dimensionality reduction techniques on a novel corpus in Spanish with electronic (online) citizen participation discussions, which represent an underexplored domain in the field.  \nConsidering a rich argument model with several argument relations, and addressing the argumenta-  \n1The majority of feature-based AM approaches consider the sentence as the argume","cbCaihNELWfSASNa","https://ap.wps.com/l/cbCaihNELWfSASNa","pdf",832412,7,1,11,"English","en",105,"# Introduction\n## Argument mining tasks and feature-based ML\n## Dimensionality reduction motivation and contributions\n# Related work\n## Feature-based machine learning for argument mining","[{\"question\":\"Which argument mining tasks are used to measure effectiveness?\",\"answer\":\"Effectiveness is evaluated on argumentative fragment detection, argument component classification, and argumentative relation recognition.\"}]","Dimensionality Reduction for Machine Learning-based Argument Mining | PDF",1785904232,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":80,"head_meta":82,"extra_data":84,"updated_unix":29},"dimensionality-reduction-for-machine-learning-based-argument-mining","",{"@graph":37,"@context":79},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/dimensionality-reduction-for-machine-learning-based-argument-mining/126279/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73],{"name":74,"@type":75,"acceptedAnswer":76},"Which argument mining tasks are used to measure effectiveness?","Question",{"text":77,"@type":78},"Effectiveness is evaluated on argumentative fragment detection, argument component classification, and argumentative relation recognition.","Answer","https://schema.org",{"og:url":53,"og:type":81,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":83,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":86},[87,91,95,99,104,109,113,116,121,124,128],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":88,"show_sort_weight":89,"slug":90},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":92,"show_sort_weight":93,"slug":94},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Exam",70,"exam",{"id":100,"doc_module":4,"doc_module_name":47,"category_name":101,"show_sort_weight":102,"slug":103},5,"Comic",60,"comic",{"id":105,"doc_module":4,"doc_module_name":47,"category_name":106,"show_sort_weight":107,"slug":108},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":110,"show_sort_weight":111,"slug":112},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":100,"slug":131},19,"General","general"]