[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120319-en":3,"doc-seo-120319-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},120319,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","Pure Component Property Estimation Framework Using Explainable Machine Learning Methods - Paper","Accurate prediction of pure component physiochemical properties supports process integration, multiscale modeling, and optimization. An enhanced framework is proposed using explainable machine learning, where a molecular representation based on a connectivity matrix generates features capturing atomic bonding. A random forest model performs feature ranking and pooling, while adjusted R² penalizes added features to quantify true contribution. Results for Tb, Lmv, Tc, and Pc using ANN and Gaussian Process Regression validate the representation, reduce test-set RMSE by up to 83.8% versus GC models, and Shapley-based analysis confirms mechanistic feature-property relationships.","\u0013 MACROBUTTON MTEditEquationSection2 公式章 1 节 1\u0013 SEQ MTEqn \\r \\h \\* MERGEFORMAT \u0015\u0013 SEQ MTSec \\r 1 \\h \\* MERGEFORMAT \u0015\u0013 SEQ MTChap \\r 1 \\h \\* MERGEFORMAT \u0015\u0015Pure Component Property Estimation Framework Using Explainable Machine Learning Methods\nJianfeng Jiao1, Xi Gao2,3,§ and Jie Li1,\n\n1Centre for Process Integration, Department of Chemical Engineering, School of Engineering, The University of Manchester, Manchester M13 9PL, UK\n2School of Electronic and Information Engineering, Tongji University, Shanghai, China 201804\n3School of Mechanical and Electrical Engineering, Jinggangshan University, Ji’an, Jiangxi, China 343009\nAbstract\nAccurate prediction of pure component physiochemical properties is crucial for process integration, multiscale modeling, and optimization. In this work, an enhanced framework for pure component property prediction by using explainable machine learning methods is proposed. In this framework, the molecular representation method based on the connectivity matrix effectively considers atomic bonding relationships to automatically generate features. The supervised machine learning model random forest is applied for feature ranking and pooling. The adjusted R² is introduced to penalize the inclusion of additional features, providing an assessment of the true contribution of features. The prediction results for normal boiling point (Tb), liquid molar volume (Lmv), critical temperature (Tc) and critical pressure (Pc) obtained using Artificial Neural Network and Gaussian Process Regression models confirm the accuracy of the molecular representation method. Comparison with GC based models shows that the root-mean-square error on the test set can be reduced by up to 83.8%. To enhance the interpretability of the model, a feature analysis method based on Shapley values is employed to determine the contribution of each feature to the property predictions. The results indicate that using the feature pooling method reduces the number of features from 13316 to 100 without compromising model accuracy. The feature analysis results for Tb, Lmv, Tc, and Pc confirms that different molecular properties are influenced by different structural features, aligning with mechanistic interpretations. In conclusion, the proposed framework is demonstrated to be feasible and provides a solid foundation for mixture component reconstruction and process integration modelling.\nKeywords: Thermodynamic properties, explainable machine learning, molecular engineering, shapley value, adjusted R².\n\n\nHighlights\n•An enhanced framework using explainable AI is proposed for property estimation\n•The connectivity matrix method is employed to generate molecular features\n•Random forest is used for feature pooling while preserving the original features\n•Adjusted R2 is used to evaluate the contribution of features to the model\n•Shapley value confirms that properties are influenced by molecule structure\n\nGraphical abstract\n\u0001\n\n\n1Introduction\nThe petroleum refining and chemical industry is the third largest greenhouse gas emitter among all stationary emission sources, accounting for nearly 5% of global greenhouse gas emissions in the energy sector \u0013ADDIN CSL_CITATION {\"citationItems\":[{\"id\":\"ITEM-1\",\"itemData\":{\"DOI\":\"10.1038/s41558-021-01072-z\",\"ISSN\":\"1758-6798\",\"abstract\":\"The role of peatlands in future climate change is uncertain because peat-derived greenhouse gas emissions are difficult to predict. Now research shows that reduced methane emissions from drying peatlands are likely to be outweighed by increasing CO2 emissions.\",\"author\":[{\"dropping-particle\":\"\",\"family\":\"Morris\",\"given\":\"Paul J\",\"non-dropping-particle\":\"\",\"parse-names\":false,\"suffix\":\"\"}],\"container-title\":\"Nature Climate Change\",\"id\":\"ITEM-1\",\"issue\":\"7\",\"issued\":{\"date-parts\":[[\"2021\"]]},\"page\":\"561-562\",\"title\":\"Wetter is better for peat carbon\",\"type\":\"article-journal\",\"volume\":\"11\"},\"uris\":[\"http://www.mendeley.com/documents/?uuid=2eebbc79-6852-493d-965b-f1bc18087e43\"]}],\"mendeley\":{\"f","cbCaihZpkQ4WBaBF","https://ap.wps.com/l/cbCaihZpkQ4WBaBF","docx",24312159,1,"English","en",105,"# Abstract\n# Highlights\n# Graphical abstract\n# 1 Introduction","[{\"question\":\"What problem does the proposed framework address?\",\"answer\":\"It addresses accurate prediction of pure component physiochemical properties, which is crucial for process integration, multiscale modeling, and optimization.\"},{\"question\":\"How are molecular features generated in the framework?\",\"answer\":\"Molecular features are automatically generated from a connectivity matrix representation that captures atomic bonding relationships.\"},{\"question\":\"How does the framework improve both accuracy and interpretability?\",\"answer\":\"Random forest is used for feature ranking and pooling, with adjusted R² penalizing extra features to assess contribution. Shapley value-based feature analysis then explains how structural features influence predicted properties.\"}]","Pure Component Property Estimation Framework Using Explainable Machine Learning Methods - Paper | DOCX",1785729437,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"pure-component-property-estimation-framework-using-explainable-machine-learning-methods-paper","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/pure-component-property-estimation-framework-using-explainable-machine-learning-methods-paper/120319/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/vnd.openxmlformats-officedocument.wordprocessingml.document","2026-08-03",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"What problem does the proposed framework address?","Question",{"text":73,"@type":74},"It addresses accurate prediction of pure component physiochemical properties, which is crucial for process integration, multiscale modeling, and optimization.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"How are molecular features generated in the framework?",{"text":78,"@type":74},"Molecular features are automatically generated from a connectivity matrix representation that captures atomic bonding relationships.",{"name":80,"@type":71,"acceptedAnswer":81},"How does the framework improve both accuracy and interpretability?",{"text":82,"@type":74},"Random forest is used for feature ranking and pooling, with adjusted R² penalizing extra features to assess contribution. Shapley value-based feature analysis then explains how structural features influence predicted properties.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]