[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117245-en":3,"doc-seo-117245-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},117245,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Text Augmentation Using a Graph-Based Approach and Clonal Selection Algorithm","Annotated data is essential for machine learning, yet producing large-scale datasets with high-quality labels remains slow and labor-intensive. Many existing NLP and ML systems depend on human annotations, whose quality varies and is constrained by arbitrary and ambiguous labeling standards, leading to unreliable supervision. This work presents an automated approach that boosts both the quality and quantity of training data for two cybersecurity tasks—fake news identification and sensitive data leakage—using CLONALG and AMR graphs, improving classifier performance by at least 5% across two datasets.","Text augmentation using a graph-based approach and clonal selection algorithm  \nHadeer Ahmed a,∗, Issa Traore a, Mohammad Mamun b, Sherif Saad c  \na ECE Department, University of Victoria, British Columbia, Canada b National Research Council Canada, New Brunswick, Canada c School of Computer Science, University of Windsor, Ontario, Canada  \n\n| A R T I C L E I N F O |  | A B S T R A C T |\n| --- | --- | --- |\n| Keywords:\u003Cbr>Data augmentation\u003Cbr>Unstructured data Cybersecurity Text generation\u003Cbr>Clonal selection |  | Annotated data is critical for machine learning models, but producing large amounts of data with high-quality labeling is a time-consuming and labor-intensive process. Natural language processing (NLP) and machine learning models have traditionally relied on the labels given by human annotators with varying degrees of competency, training, and experience. These kinds of labels are incredibly problematic because they are defined and enforced by arbitrary and ambiguous standards. In order to solve these issues of insufficient high-quality labels, researchers are now investigating automated methods for enhancing training and testing data sets. In this paper, we demonstrate how our proposed method improves the quality and quantity of data in two cybersecurity problems (fake news identification & sensitive data leak) by employing the clonal selection algorithm (CLONALG) and abstract meaning representation (AMR) graphs, and how it improves the performance of a classifier by at least 5% on two datasets. |\n\n1. Introduction  \nThe recent advances in machine learning (ML) and deep learning (DL) gave rise to powerful prediction models with applications in many sectors. Machine learning and deep learning models, in particular, are data-driven algorithms that rely heavily on data to solve problems. Instead of using explicit rules to solve problems, ML and DL models learn to solve problems by analyzing data (aka ‘‘training data’’) in a process known as model training. Therefore, the availability of training data is crucial for designing and developing reliable systems that utilize machine learning and deep learning models. Data availability is a wellknown problem in machine learning; the unavailability of training data is a common challenge facing data scientists when developing ML or DL models. Moreover, some sectors are known to struggle to collect enough training data. For example, in healthcare, finance, and cybersecurity, many technical and business constraints prevent organizations from obtaining enough training data to build reliable machine learning models. Additionally, the datasets available are quite often unreliable. Not only are these datasets used for training, but they are also used to test and evaluate models to assure their functionality and establish their overall efficacy. Hence, the quality of the data has a significant effect on the quality of the models. It is possible to train machine learning models using unreliable data. But doing so could lead to predictions that do not match reality because of biases or mistakes. In some instances, even  \nwhen the data originates from a reliable source, it is unsafe to depend on data that has not been sufficiently verified.  \nIn general, machine learning models that have been verified or trained against well-known benchmarks are considered cutting-edge for the problem they were created to solve. But because there is not enough data, these benchmarks are frequently used to train various machine learning models that address different, non-identical tasks. When this occurs, the datasets utilized are unlikely to be an accurate reflection of the data that some of those models would be applied to in the real world, leading to incorrect predictions by the models (Joshi, 2021).  \nThere are also problems with incorrect labels, which are the annotations that many models utilize to identify correlations in data. A significant number of these labels are assigned by human operators in","cbCaikEDdOCHlAye","https://ap.wps.com/l/cbCaikEDdOCHlAye","pdf",1103679,1,"English","en",105,"# Introduction\n## Data availability challenges\n## Label quality issues in manual annotation\n## Related solutions and motivation\n# Method and contributions","[{\"question\":\"Why is data augmentation important in machine learning for cybersecurity tasks?\",\"answer\":\"Model quality depends heavily on training and testing data. When training data is scarce or unreliable, augmentation methods help improve the dataset used for learning and evaluation.\"},{\"question\":\"What problems does the proposed approach address regarding labels?\",\"answer\":\"The paper highlights incorrect or inconsistent labels caused by subjective, biased, or low-quality human annotation, which can distort learned correlations and lead to unreliable predictions.\"},{\"question\":\"How does the method use CLONALG and AMR graphs to improve results?\",\"answer\":\"It employs the clonal selection algorithm together with abstract meaning representation (AMR) graphs to enhance both the quantity and quality of training data, yielding at least a 5% improvement in classifier performance on two datasets.\"}]","Text Augmentation Using a Graph-Based Approach and Clonal Selection Algorithm | PDF",1785674642,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"text-augmentation-using-a-graph-based-approach-and-clonal-selection-algorithm","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/text-augmentation-using-a-graph-based-approach-and-clonal-selection-algorithm/117245/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is data augmentation important in machine learning for cybersecurity tasks?","Question",{"text":74,"@type":75},"Model quality depends heavily on training and testing data. When training data is scarce or unreliable, augmentation methods help improve the dataset used for learning and evaluation.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What problems does the proposed approach address regarding labels?",{"text":79,"@type":75},"The paper highlights incorrect or inconsistent labels caused by subjective, biased, or low-quality human annotation, which can distort learned correlations and lead to unreliable predictions.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the method use CLONALG and AMR graphs to improve results?",{"text":83,"@type":75},"It employs the clonal selection algorithm together with abstract meaning representation (AMR) graphs to enhance both the quantity and quality of training data, yielding at least a 5% improvement in classifier performance on two datasets.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]