[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125861-en":3,"doc-seo-125861-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125861,1099523882367,"Hazel","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Data augmentation for machine learning of chemical process flowsheets","Artificial intelligence can accelerate the design and engineering of chemical processes, but it depends on large datasets while machine-readable flowsheet data remains limited. Prior work showed transformer language models for flowsheet autocompletion using SFILES 2.0 strings and translation from PFDs to P&IDs. This study introduces a SFILES 2.0-based data augmentation methodology. Results on an autocompletion task improve prediction uncertainty by 14.7%, enabling broader use with other SFILES-driven machine learning algorithms.","Data augmentation for machine learning of chemical process flowsheets  \nLukas Schulze Balhorna, Edwin Hirtreitera, Lynn Luderera, Artur M.  \nSchweidtmanna,*  \na Process Intelligence Research, Department of Chemical Engineering, Delft University of Technology, Van der Maasweg 9, Delft 2629 HZ, The Netherlands  \n*Corresponding author. Email: [a.schweidtmann@tudelft.nl](a.schweidtmann@tudelft.nl)  \nAbstract  \nArtificial intelligence has great potential for accelerating the design and engineering of chemical processes. Recently, we have shown that transformer-based language models can learn to auto-complete chemical process flowsheets using the SFILES 2.0 string notation. Also, we showed that language translation models can be used to translate Process Flow Diagrams (PFDs) into Process and Instrumentation Diagrams (P&IDs) . However, artificial intelligence methods require big data and flowsheet data is currently limited. To mitigate this challenge of limited data, we propose a new data augmentation methodology for flowsheet data that is represented in the SFILES 2.0 notation. We show that the proposed data augmentation improves the performance of artificial intelligencebased process design models. In our case study flowsheet data augmentation improved the prediction uncertainty of the flowsheet autocompletion model by 14.7%. In the future, our flowsheet data augmentation can be used for other machine learning algorithms on chemical process flowsheets that are based on SFILES notation.  \nKeywords: Data Augmentation, Flowsheet Autocompletion, SFILES, Transformers  \n1. Introduction  \nThe design of a flowsheet topology is an important step in early process synthesis. This step consists of selecting and arranging unit operations for a chemical process. Artificial intelligence (AI) methods have the potential to learn from previous flowsheets and support engineers in process development (Hirtreiter et al., 2022; Oeing et al., 2022; Schweidtmann, 2022; Vogel et al., 2023) . For instance, Vogel et al. (2023) proposed an algorithm for the autocompletion of flowsheets. This autocompletion algorithm is inspired by text-autocompletion from natural language processing (NLP) that is based on generative transformer models (Radford et al., 2019) . In addition, Hirtreiter et al. (2022) showed that the prediction of control structure elements from Process Flow Diagrams (PFDs) can be interpreted as a translation task between PFDs and Process and Instrumentation Diagrams (P&IDs) . Hence, they deployed a sequence-to-sequence transformer architecture which is commonly used for translation of text between different languages. These flowsheet transformers rely on machine-readable flowsheet representations.  \nTo represent flowsheets in a machine-readable format, we depict them as graphs or as text using unique, i.e., canonical, SFILES 2.0 strings (Vogel et al., 2022b) . In general, flowsheets are drawings of chemical processes. Chemical engineers use flowsheets for the communication, planning, operation, simulation, and construction of these processes. An example flowsheet is given in Figure 1.  \nAn intuitive way to represent flowsheets is via graphs with unit operations as nodes and stream connections as directed edges. Besides the graph representation, flowsheets can also be represented as strings. D’Anterroches (2005) introduced the Simplified Flowsheet Input-Line Entry-System (SFILES) notation, which we recently extended to include control structures and other features in (Vogel et al., 2022b) . When creating the SFILES, we traverse the graph by starting at an input node and following the stream direction until we reach a product node or a recycle. In case the stream branches at a node, i.e., a splitter, we need to decide which stream to follow first. To determine the order of the branches in the linear string, the SFILES algorithm assigns each node a unique rank. The SFILES string for the flowsheet from Figure 1 is given by:  \n(raw)(hex){1}(r)\u003C&|(raw)","cbCaikwnctyeDAsv","https://ap.wps.com/l/cbCaikwnctyeDAsv","pdf",251803,1,6,"English","en",105,"# Introduction\n## Flowsheet representation and SFILES 2.0\n## Data limitation and need for augmentation\n## Proposed augmentation approach","[{\"question\":\"Why is data augmentation needed for chemical process flowsheet machine learning?\",\"answer\":\"Flowsheet data is often limited because most flowsheets are stored as images and are not machine-readable, and publicly available machine-readable datasets are scarce.\"},{\"question\":\"What representation does the proposed method use for flowsheets?\",\"answer\":\"The method augments flowsheet data represented as canonical SFILES 2.0 strings.\"},{\"question\":\"How does the augmentation improve model performance in the study?\",\"answer\":\"In the case study, flowsheet data augmentation improves the prediction uncertainty of the flowsheet autocompletion model by 14.7%.\"}]","Data augmentation for machine learning of chemical process flowsheets | PDF",1785901638,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"data-augmentation-for-machine-learning-of-chemical-process-flowsheets","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/data-augmentation-for-machine-learning-of-chemical-process-flowsheets/125861/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is data augmentation needed for chemical process flowsheet machine learning?","Question",{"text":75,"@type":76},"Flowsheet data is often limited because most flowsheets are stored as images and are not machine-readable, and publicly available machine-readable datasets are scarce.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What representation does the proposed method use for flowsheets?",{"text":80,"@type":76},"The method augments flowsheet data represented as canonical SFILES 2.0 strings.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the augmentation improve model performance in the study?",{"text":84,"@type":76},"In the case study, flowsheet data augmentation improves the prediction uncertainty of the flowsheet autocompletion model by 14.7%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]