[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123647-en":3,"doc-seo-123647-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123647,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Automatic recognition of complementary strands: lessons regarding machine learning abilities in RNA folding","The study addresses limitations in using machine learning to predict RNA secondary structure from single sequences, focusing on why overfitting restricts learning of core folding mechanisms. A related binary task is analyzed: whether two sequences are fully complementary. Results compare how model architecture, capacity, and dataset size and composition affect classification accuracy. Low-capacity models better handle mislabelled training data, whereas large capacities improve generalization to structurally dissimilar cases, and neural networks struggle with base-complementarity concepts in lengthwise extrapolation.","TYPE Original Research PUBLISHED 04 September 2023 DOI 10.3389/fgene.2023.1254226  \nOPEN ACCESS  \nEDITED BY  \nYadong Zheng,  \nZhejiang Agriculture and Forestry University, China  \nREVIEWED BY  \nJianhua Jia,  \nJingdezhen Ceramic Institute, China Cuncong Zhong,  \nUniversity of Kansas, United States  \n*CORRESPONDENCE  \nFrançois Major,  \n [francois.major@umontreal.ca](francois.major@umontreal.ca)  \nRECEIVED 06 July 2023  \nACCEPTED 16 August 2023  \nPUBLISHED 04 September 2023  \nCITATION  \nChasles S and Major F (2023), Automatic recognition of complementary strands: lessons regarding machine learning abilities in RNA folding.  \nFront. Genet. 14:1254226 .  \ndoi: 10.3389/fgene.2023.1254226  \nCOPYRIGHT  \n© 2023 Chasles and Major. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nAutomatic recognition of complementary strands: lessons regarding machine learning abilities in RNA folding  \nSimon Chasles 1,2 and François Major 1,2*  \n1Institute for Research in Immunology and Cancer, Montréal, QC, Canada, 2Department of Computer Science and Operations Research, Université de Montréal, Montréal, QC, Canada  \nIntroduction: Prediction of RNA secondary structure from single sequences still needs substantial improvements. The application of machine learning (ML) to this problem has become increasingly popular. However, ML algorithms are prone to overﬁtting, limiting the ability to learn more about the inherent mechanisms governing RNA folding. It is natural to use high-capacity models when solving such a difﬁcult task, but poor generalization is expected when too few examples are available.  \nMethods: Here, we report the relation between capacity and performance on a fundamental related problem: determining whether two sequences are fully complementary. Our analysis focused on the impact of model architecture and capacity as well as dataset size and nature on classiﬁcation accuracy.  \nResults: We observed that low-capacity models are better suited for learning with mislabelled training examples, while large capacities improve the ability to generalize to structurally dissimilar data. It turns out that neural networks struggle to grasp the fundamental concept of base complementarity, especially in lengthwise extrapolation context.  \nDiscussion: Given a more complex task like RNA folding, it comes as no surprise that the scarcity of useable examples hurdles the applicability of machine learning techniques to this ﬁeld.  \nKEYWORDS  \nRNA folding, base complementarity, machine learning, neural networks, binary classiﬁcation, artiﬁcial data  \nIntroduction  \nIdentifying potential structural candidates for a single RNA sequence is a computationally demanding task. The Zuker-style dynamic programming approach to fold an RNA sequence of length L without pseudoknots requires time complexity in O (L3 ) (Zuker and Stiegler 1981; Hofacker et al., 1994) . Algorithms that take into account pseudoknots are even more complex and have been reported to require signiﬁcantly more computational power ranging from O (L4 ) to O (L6 ) (Rivas and Eddy 1999; Condon et al., 2004), or higher (Marchand et al., 2022) .  \nMachine learning (ML) algorithms offer an alternative to traditional methods for identifying RNA structural candidates. In particular, neural networks can compute structures in an end-to-end fashion, allowing for quick inference in a single feedforward  \nFrontiers in Genetics 01 [frontiersin.org](frontiersin.org)  \nChasles and Major 10.3389/fgene.2023.1254226  \nFIGURE 1  \nThe four neural network architectures. The 4 tested neural network architectures take nucleotid","cbCaieU9x0SzdfZh","https://ap.wps.com/l/cbCaieU9x0SzdfZh","pdf",2997345,1,12,"English","en",105,"# Introduction\n## RNA structure prediction and computational complexity\n## Machine learning for RNA tasks\n# Methods\n## Capacity-performance analysis for sequence complementarity\n## Model architecture and dataset effects\n# Results\n## Low vs high model capacity under mislabelling\n## Generalization to structurally dissimilar data\n## Difficulties in base complementarity learning\n# Discussion\n## Example scarcity and ML applicability","[{\"question\":\"What problem does the paper focus on beyond RNA folding prediction?\",\"answer\":\"It studies a fundamental related classification task: determining whether two sequences are fully complementary. This helps analyze how model capacity and training data influence learning behavior.\"},{\"question\":\"How do model capacity levels affect performance and generalization?\",\"answer\":\"Low-capacity models perform better when training examples contain mislabelling, while larger capacities improve generalization to structurally dissimilar data.\"},{\"question\":\"What challenge do neural networks face according to the results?\",\"answer\":\"Neural networks struggle to grasp the concept of base complementarity, particularly in lengthwise extrapolation settings.\"}]","Automatic recognition of complementary strands: lessons regarding machine learning abilities in RNA folding | PDF",1785817823,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"automatic-recognition-of-complementary-strands-lessons-regarding-machine-learning-abilities-in-rna-folding","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/automatic-recognition-of-complementary-strands-lessons-regarding-machine-learning-abilities-in-rna-folding/123647/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper focus on beyond RNA folding prediction?","Question",{"text":75,"@type":76},"It studies a fundamental related classification task: determining whether two sequences are fully complementary. This helps analyze how model capacity and training data influence learning behavior.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do model capacity levels affect performance and generalization?",{"text":80,"@type":76},"Low-capacity models perform better when training examples contain mislabelling, while larger capacities improve generalization to structurally dissimilar data.",{"name":82,"@type":73,"acceptedAnswer":83},"What challenge do neural networks face according to the results?",{"text":84,"@type":76},"Neural networks struggle to grasp the concept of base complementarity, particularly in lengthwise extrapolation settings.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]