[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126025-en":3,"doc-seo-126025-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126025,2336474466412,"Ezra","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","On the use of machine learning to generate in-silico data for batch process monitoring under small-data scenarios - Research and analysis","Batch process monitoring with PCA typically requires sufficient historical manufacturing trajectories to learn normal operating conditions, which fails in small-data cases common when a facility produces a new product. This work enhances a literature data-driven approach by using machine learning based on Gaussian process state-space models to generate in-silico batch trajectories from limited history, then combines real and synthetic data to build monitoring models. Automatic parameter tuning, plus monitoring-oriented indicators, support scalable industrial deployment and validation on benchmark penicillin semi-batch datasets.","Computers and Chemical Engineering 180 (2024) 108469  \nContents lists available at ScienceDirect  \nComputers and Chemical Engineering  \njournal [homepage:](homepage: www.elsevier.com/locate/compchemeng)[ www.elsevier.com/locate/compchemeng](homepage: www.elsevier.com/locate/compchemeng)  \n| On the use of machine learning to generate in-silico data for batch process   monitoring under small-data scenarios\u003Cbr>Luca Gasparini a, b, Antonio Benedetti c, Giulia Marchesed, Connor Gallagher e,\u003Cbr>Pierantonio Facco *\u003Cbr>a, Massimiliano Barolo a,\u003Cbr>a CAPE-Lab – Computer-Aided Process Engineering Laboratory, Department of Industrial Engineering, University of Padova, via Marzolo 9, 35131 Padova PD, Italy b INSTM – Consorzio Interuniversitario Nazionale per la Scienza e la Tecnologia dei Materiali, via Giusti 9, 50121 Firenze FI, Italy\u003Cbr>c Process Engineering & Analytics, Medicine Development and Supply, GSK R&D, Park Rd Ware, SG12 0DP, United Kingdom\u003Cbr>d Process Engineering & Analytics, Medicine Development and Supply, GSK R&D, Gunnels Wood Rd, Stevenage, SG1 2NY, United Kingdome Process Engineering & Analytics, Medicine Development and Supply, GSK R&D, 1250 S. Collegeville Road, Collegeville, PA 19426, United States |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |\n| Keywords:\u003Cbr>Batch processes Small data Big data Machine learning Process monitoring\u003Cbr>Biopharmaceutical industry Pharmaceutical engineering |  | Batch process monitoring using principal component analysis requires sufficient historical manufacturing data to model the normal operating conditions of the process. However, when a new product is to be manufactured for the first time in a given facility, very limited historical data are available, thus entailing a small-data scenario. We thoroughly investigate and improve a data-driven methodology, previously reported in the literature (Tulsyan, Garvin & Ündey (2019). J. Process Control, 77, 114–133), that enables batch process monitoring under such typeof scenarios. The methodology exploits machine learning algorithms (based on Gaussian process state-space models) to generate in-silico batch trajectory data from the few available historical ones, and then uses the overall pool of real and in-silico data to build a process monitoring model. We develop automatic procedures to tune the values of several parameters of this machine-learning framework, in such a way that the generation of consistent in-silico batch trajectory data can be streamlined, thus facilitating the deployment of the framework atan industrial level. Furthermore, we develop indicators and a metric to assist the in-silico data generation activity from a process monitoring-relevant perspective. Finally, using datasets from a benchmark simulated semi-batch process for the manufacturing of penicillin, we thoroughly investigate the appropriateness of the in-silico generated data for the purpose of process monitoring. |\n\n1. Introduction  \nIn batch and semi-batch manufacturing, reproducibility (or consistency) across batches is required to guarantee that the end-product quality targets are met after every batch. Whether or not a new batch conforms to a set of “normal” batches that were run in the past can be assessed by using a data-driven process monitoring framework, where the most widely used one exploits multivariate statistical techniques, such as principal component analysis (PCA; (Jackson, 1991; Wise and Gallagher, 1996; Kourti, 2003)). The rationale behind this method is quite simple: (a) collect a set of historical batches that were run satisfactorily (“normal” batches); (b) build a PCA model on the trajectories of all measured variables across all normal batches; (c) using statistical control charts built through the PCA model, test whether a new batch can be considered normal; (d) if it is not, raise an alarm (fault diagnosis  \nmay then follow). This approach is very effective (Kourti et al., 1995; Reis and Gins, 2017), but suffe","cbCaigsfogiiq1gC","https://ap.wps.com/l/cbCaigsfogiiq1gC","pdf",9036995,5,1,26,"English","en",105,"# Introduction\n## Background: batch and semi-batch monitoring\n## PCA-based process monitoring and the low-N limitation\n# Proposed methodology and small-data solution\n## In-silico trajectory generation via Gaussian process state-space models\n## Parameter tuning and deployment readiness\n## Monitoring indicators and validation","[{\"question\":\"Why is principal component analysis challenging under small-data batch monitoring scenarios?\",\"answer\":\"PCA-based monitoring relies on identifying normal operating conditions from historical trajectories, which typically requires many historical batches. In small-data (low-N) cases, available history is too limited to build reliable PCA models.\"},{\"question\":\"How does the proposed method generate in-silico data when historical data are scarce?\",\"answer\":\"It uses machine learning algorithms based on Gaussian process state-space models to create in-silico batch trajectory data from the few available historical trajectories.\"},{\"question\":\"What is the role of combining real and in-silico data in the monitoring model?\",\"answer\":\"The approach builds the process monitoring model using the overall pool of real and generated in-silico data, improving model construction when historical information is insufficient.\"}]","On the use of machine learning to generate in-silico data for batch process monitoring under small-data scenarios - Research and analysis | PDF",1785902610,66,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"on-the-use-of-machine-learning-to-generate-in-silico-data-for-batch-process-monitoring-under-small-data-scenarios-research-and-analysis","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/on-the-use-of-machine-learning-to-generate-in-silico-data-for-batch-process-monitoring-under-small-data-scenarios-research-and-analysis/126025/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is principal component analysis challenging under small-data batch monitoring scenarios?","Question",{"text":77,"@type":78},"PCA-based monitoring relies on identifying normal operating conditions from historical trajectories, which typically requires many historical batches. In small-data (low-N) cases, available history is too limited to build reliable PCA models.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the proposed method generate in-silico data when historical data are scarce?",{"text":82,"@type":78},"It uses machine learning algorithms based on Gaussian process state-space models to create in-silico batch trajectory data from the few available historical trajectories.",{"name":84,"@type":75,"acceptedAnswer":85},"What is the role of combining real and in-silico data in the monitoring model?",{"text":86,"@type":78},"The approach builds the process monitoring model using the overall pool of real and generated in-silico data, improving model construction when historical information is insufficient.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]