[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118625-en":3,"doc-seo-118625-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118625,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","MLRegTest - A Benchmark for the Machine Learning of Regular Languages","Synthetic datasets constructed from formal languages enable fine-grained study of how machine learning systems learn and generalize in sequence classification. The paper introduces MLRegTest, a benchmark containing training, development, and test sets spanning 1,800 regular languages. It groups languages by logical complexity and by the types of logical literals used, offering a systematic lens on long-distance dependencies. Experiments evaluate neural models including simple RNN, LSTM, GRU, and transformers, showing that performance varies with test set design, language class, and architecture.","arXiv :2304 .07687v4 [ cs .LG] 1 Sep 2024  \nMLRegTest: A Benchmark for the Machine Learning of  \nRegular Languages  \nSam van der Poel  \nSchool of Mathematics  \nGeorgia Institute of Technology  \nDakotah Lambert  \nDepartment of Computer Science Haverford College  \nKalina Kostyszyn  \nDepartment of Linguistics &  \nInstitute of Advanced Computational Science Stony Brook University  \nTiantian Gao Rahul Verma  \nDepartment of Computer Science Stony Brook University  \nDerek Andersen Joanne Chau Emily Peterson Cody St. Clair  \nDepartment of Linguistics Stony Brook University  \nPaul Fodor  \nDepartment of Computer Science Stony Brook University  \nChihiro Shibata  \nDepartment of Advanced Sciences Graduate School of Science and Engineering Hosei University  \nJeffrey Heinz  \nDepartment of Linguistics &  \nInstitute of Advanced Computational Science Stony Brook University  \n[samvanderpoel@gatech.edu](samvanderpoel@gatech.edu)  \n[dakotahlambert@acm.org](dakotahlambert@acm.org)  \n[kalina.kostyszyn@stonybrook.edu](kalina.kostyszyn@stonybrook.edu)  \n[tiagao@cs.stonybrook.edu](tiagao@cs.stonybrook.edu)[rxverma1@gmail.com](rxverma1@gmail.com)  \n[derek.andersen@alumni.stonybrook.edu](derek.andersen@alumni.stonybrook.edu)[ ](derek.andersen@alumni.stonybrook.edu)[choryan.chau@alumni.stonybrook.edu](choryan.chau@alumni.stonybrook.edu)[ ](choryan.chau@alumni.stonybrook.edu)[emily.peterson@alumni.stonybrook.edu](emily.peterson@alumni.stonybrook.edu)[ ](emily.peterson@alumni.stonybrook.edu)[cody.stclair@alumni.stonybrook.edu](cody.stclair@alumni.stonybrook.edu)  \n[pfodor@cs.stonybrook.edu](pfodor@cs.stonybrook.edu)  \n[chihiro@hosei.ac.jp](chihiro@hosei.ac.jp)  \n[jeffrey.heinz@stonybrook.edu](jeffrey.heinz@stonybrook.edu)  \n©2024 Sam van der Poel, Dakotah Lambert, Kalina Kostyszyn, Tiantian Gao, Rahul Verma, Derek Andersen, Joanne Chau, Emily Peterson, Cody St. Clair, Paul Fodor, Chihiro Shibata, and Jeffrey Heinz.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/) .  \nvan der Poel et al.  \nAbstract  \nSynthetic datasets constructed from formal languages allow fine-grained examination of the learning and generalization capabilities of machine learning systems for sequence classification. This article presents a new benchmark for machine learning systems on sequence classification called MLRegTest, which contains training, development, and test sets from  \n1,800 regular languages.  \nDifferent kinds of formal languages represent different kinds of long-distance dependencies, and correctly identifying long-distance dependencies in sequences is a known challenge for ML systems to generalize successfully. MLRegTest organizes its languages according to their logical complexity (monadic second order, first order, propositional, or restricted propositional) and the kind of logical literals (string, tier-string, subsequence, or combinations thereof) . The logical complexity and choice of literal provides a systematic way to understand different kinds of long-distance dependencies in regular languages, and therefore to understand the capacities of different ML systems to learn such long-distance dependencies.  \nFinally, the performance of different neural networks (simple RNN, LSTM, GRU, transformer) on MLRegTest is examined. The main conclusion is that performance depends significantly on the kind of test set, the class of language, and the neural network architecture.  \nKeywords: formal languages, regular languages, subregular languages, sequence classification, neural networks, long-distance dependencies  \n1 Introduction  \nThis article presents a new benchmark for the machine learning (ML) of regular languages called MLRegTest.1 Regular languages are formal languages, which are sets of sequences definable with certain kinds of formal grammars, including regular expressions, finite-state acceptors, and monadic second order logic with either the successor or precedence relation in the model signature","cbCaiqZG4tCbnxGF","https://ap.wps.com/l/cbCaiqZG4tCbnxGF","pdf",2189799,1,45,"English","en",105,"# Introduction\n# MLRegTest benchmark description\n## Language organization by logical complexity\n## Literals and long-distance dependencies\n# Experimental evaluation\n## Neural network architectures\n## Results analysis","[{\"question\":\"What problem does MLRegTest address in machine learning for regular languages?\",\"answer\":\"MLRegTest targets the challenge of understanding and diagnosing how well ML systems learn long-distance dependencies needed for successful generalization in sequence classification.\"},{\"question\":\"How is the MLRegTest benchmark structured?\",\"answer\":\"It provides training, development, and test sets across 1,800 regular languages, with languages organized by logical complexity and the kinds of logical literals used.\"},{\"question\":\"Which neural network models are evaluated on MLRegTest and what influences performance?\",\"answer\":\"The study examines simple RNN, LSTM, GRU, and transformer models. Performance depends strongly on the test set type, the class of language, and the neural network architecture.\"}]","MLRegTest - A Benchmark for the Machine Learning of Regular Languages | PDF",1785684578,113,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mlregtest-a-benchmark-for-the-machine-learning-of-regular-languages","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mlregtest-a-benchmark-for-the-machine-learning-of-regular-languages/118625/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MLRegTest address in machine learning for regular languages?","Question",{"text":75,"@type":76},"MLRegTest targets the challenge of understanding and diagnosing how well ML systems learn long-distance dependencies needed for successful generalization in sequence classification.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the MLRegTest benchmark structured?",{"text":80,"@type":76},"It provides training, development, and test sets across 1,800 regular languages, with languages organized by logical complexity and the kinds of logical literals used.",{"name":82,"@type":73,"acceptedAnswer":83},"Which neural network models are evaluated on MLRegTest and what influences performance?",{"text":84,"@type":76},"The study examines simple RNN, LSTM, GRU, and transformer models. Performance depends strongly on the test set type, the class of language, and the neural network architecture.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]