[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119870-en":3,"doc-seo-119870-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119870,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","ML-Peaks - CHIP-Seq peak detection pipeline using machine learning techniques","CHIP-Seq data is essential for locating protein-DNA binding regions that reveal disease molecular mechanisms and guide therapeutic target discovery. Peak detection remains difficult because existing tools often rely on manual visualization inspection and can struggle with high background noise, low signal-to-noise ratio, and variable peak shapes. The proposed pipeline uses sliding-window data preprocessing and feature reduction to enable machine learning models to classify Peaks vs NoPeaks without human intervention. Experiments on an H3K9me3 TDH BP CHIP-Seq dataset achieve an F1-score of 0.9644 and a false positive rate of 0.1030.","ML-Peaks: CHIP-Seq peak detection pipeline using  \nmachine learning techniques  \nSajad Amouei Sheshkal 1 ,2 ,4 , Michael Alexander Riegler1 ,2 ,3 , Hugo Lewi Hammer1 ,2  \n1 Oslo Metropolitan University (OsloMet), Oslo, Norway  \n2 Simula Metropolitan Center for Digital Engineering, Oslo, Norway  \n3 UiT The Arctic University of Norway, Tromsø, Norway  \n4 Ifocus eye clinic, Haugesund, Norway  \nAbstract—CHIP-Seq data is critical for identifying the locations where proteins bind to DNA, offering valuable insights into disease molecular mechanisms and potential therapeutic targets. However, identifying regions of protein binding, or peaks, in CHIP-seq data can be challenging due to limitations in peak detection methods. Current computational tools often require manual human inspection using data visualization, making it challenging and resource demanding to detect all peaks, particularly in large datasets. CHIP-seq data poses difficulties in detecting peaks due to its high background noise, low signal-tonoise ratio, and variation in the size and shape of the peaks. To overcome these challenges, we propose a data preprocessing approach using sliding window and feature reduction techniques, and the resulting features can be further used in machine learning methods. Our machine learning methodology can accurately identify peaks using a small training set, which represents a distinct advantage over commonly used statistical approaches, as it has a greater capacity for learning from data.  \nWe tested our methodology on the H3K9me3 TDH BP CHIPSeq dataset exploring a range of different machine learning methods, sliding window settings, and feature reduction techniques to detect peak values without human intervention. Our pipeline efficiently detected the peaks, and achieved an F1-score of 0.9644 and a false positive rate of 0.1030.  \nIndex Terms—Peak detection, CHIP-Seq dataset, Machine learning, Sliding window, Feature reduction.  \nI. INTRODUCTION  \nChromatin immunoprecipitation followed by highthroughput sequencing (CHIP-Seq) is a widely used technique for studying the binding of DNA-associated proteins, such as transcription factors, to the genome [1] . The method enables researchers to map the binding sites of DNA-associated proteins and their target genes, providing important insights into the regulation of gene expression. The CHIP-Seq data generated from this technique requires further analysis to identify the regions of the genome where DNA-associated proteins bind. To aid in this process, researchers use a variety of software tools, including both graphical tools and peak detection algorithms. This process is known as peak detection, and it is a key step in the CHIP-Seq data analysis pipeline [2] .  \nGraphical tools, such as the UCSC genome browser [3], provide a visual interface for exploring genomic data, including the locations of functional elements. These tools allow scientists to view their data in the context of the genome,  \nalongside other relevant information, such as gene annotations and conservation information. This visual representation can be particularly useful in the early stages of data exploration, as it helps to identify trends and patterns that may not be immediately apparent from the raw data.  \nHowever, graphical tools suffer from three main disadvantages. At a single base resolution, peak start and end locations are not obvious on visual inspection. Second, visual interpretation is inherently subjective, which makes it difficult for other researchers to reproduce it. Finally, it is not enough time for researchers to visually inspect and identify peaks across the entire genome. For these reasons, it is useful to use computational methods, and peak detection algorithms, to systematically and accurately identify CHIP-Seq peaks, in addition to using the UCSC Genome Browser [3] for data visualization and exploration.  \nPeak detection algorithms are used to identify the regions of the genome that are enriched wit","cbCaioM8dKfml1Zh","https://ap.wps.com/l/cbCaioM8dKfml1Zh","pdf",1511441,1,6,"English","en",105,"# Introduction\n## Background on CHIP-Seq and peak detection\n## Limitations of graphical and statistical methods\n## Machine learning approaches for peak detection\n# Proposed ML-Peaks pipeline\n## Sliding window feature extraction\n## Dimensionality reduction and classification","[{\"question\":\"Why is CHIP-Seq peak detection challenging?\",\"answer\":\"Peak detection is difficult due to high background noise, low signal-to-noise ratio, and variability in peak size and shape. Many tools also require manual inspection, which is time-consuming for large datasets.\"},{\"question\":\"What does the ML-Peaks pipeline do?\",\"answer\":\"It uses a sliding window to extract sub-sequences, then applies feature reduction and dimensionality reduction. The resulting representations are used by machine learning models to classify regions as Peaks or NoPeaks.\"},{\"question\":\"How is the proposed approach evaluated and with what results?\",\"answer\":\"The methodology is tested on an H3K9me3 TDH BP CHIP-Seq dataset across different ML methods and preprocessing settings. The pipeline detects peaks without human intervention, reaching an F1-score of 0.9644 and a false positive rate of 0.1030.\"}]","ML-Peaks - CHIP-Seq peak detection pipeline using machine learning techniques | PDF",1785726749,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ml-peaks-chip-seq-peak-detection-pipeline-using-machine-learning-techniques","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ml-peaks-chip-seq-peak-detection-pipeline-using-machine-learning-techniques/119870/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is CHIP-Seq peak detection challenging?","Question",{"text":75,"@type":76},"Peak detection is difficult due to high background noise, low signal-to-noise ratio, and variability in peak size and shape. Many tools also require manual inspection, which is time-consuming for large datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the ML-Peaks pipeline do?",{"text":80,"@type":76},"It uses a sliding window to extract sub-sequences, then applies feature reduction and dimensionality reduction. The resulting representations are used by machine learning models to classify regions as Peaks or NoPeaks.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed approach evaluated and with what results?",{"text":84,"@type":76},"The methodology is tested on an H3K9me3 TDH BP CHIP-Seq dataset across different ML methods and preprocessing settings. The pipeline detects peaks without human intervention, reaching an F1-score of 0.9644 and a false positive rate of 0.1030.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]