[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124717-en":3,"doc-seo-124717-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124717,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Prediction of Spontaneous Preterm Birth Using Supervised Machine Learning on Metabolomic Data - Supplementary Methods","Supplementary methods detail how metabolomic measurements were processed and used to predict spontaneous preterm birth (sPTB) from cohort samples. Metabolite values were log-transformed and low-variance features removed, then six supervised classification algorithms were trained with repeated 10-fold cross-validation. A weighted Cox proportional-hazards framework modeled time-to-sPTB under a case-cohort design, and risk groups were evaluated with C-index and permutation log-rank tests. Feature selection and internal validation were performed using variable-importance ranking, elastic-net penalized logistic regression, best-subsets selection, and optimism-corrected ROC performance using Stata and R.","Supporting Information  \nContents  \nSupplementary Methods  \nTable S1. Characteristics of the study participants by birth outcome.  \nTable S2. Important metabolites identified by at least two models trained on metabolite data (n = 47) .  \nTable S3. Best subsets selection results for the optimal 1-10 predictor models.  \nTable S4 . Associations with spontaneous birth at any gestational age at term, split into three categories:  \n37 to \u003C39, 39 to \u003C41 and ≥41 weeks of gestational age.  \nFigure S1. Kaplan-Meier curves resulting from Cox-PH model trained on metabolite data with samples divided into four quartiles of estimated risk.  \nFigure S2. Kaplan-Meier curves resulting from Cox-PH model trained on metabolite data with samples dichotomized into 2 risk groups.  \nFigure S3. Heatmap visualizing correlations between the 47 metabolites identified as important by at least two models.  \nFigure S4. The means of the z scores (95% CI) of the metabolites included in the 4-predictor model at each gestational age by sPTB status.  \nFigure S5. The means of the z scores (95% CI) of the metabolites included in the 4-predictor model at each gestational age by sETB status.  \nSupplementary References  \nSupplementary Methods  \nSoftware  \nMetabolite identification analyses were performed using R, version 4.1.1. The primary software used in this analysis was the Lilikoi v2 .0 R package (lilikoi), a personalised pathway-based package for diagnosis and prognosis predictions using metabolomics data. 1,2 Changes were made to the package source code as required. The code used for the analysis has been made available at [https://github.com/lilcamwheat/thesis.git](https://github.com/lilcamwheat/thesis.git. Internal metabolite validation)[. Internal metabolite validation](https://github.com/lilcamwheat/thesis.git. Internal metabolite validation) analyses were performed using Stata version 17.0 (StataCorp, Texas, USA) .  \nClassification models  \nDue to skewness, the metabolite values (MoMs) were first log-transformed.3-5 Twenty-five metabolites were excluded for having zero or near zero variance in the study population, leaving 812 metabolites for analysis. We applied 6 algorithms to classify spontaneous preterm birth (sPTB) and control samples:  \ngeneralised boosted model (GBM), linear discriminant analysis (LDA), penalized logistic regression (LOG), random forest (RF), recursive partitioning and regression analysis (Rpart), and support vector machine (SVM) . These methods, available in the lilikoi, have been widely used in the metabolomics literature.6,7 To mitigate the risk of overfitting, a 10-fold cross-validation (CV) with 100 repeats was applied, and the averaged area under the receiver operating characteristic curve (AUC) was reported. All indicated preterm births (iPTBs) were excluded from the classification models.  \nPrognosis prediction  \nThe weighted Cox proportional-hazards (Cox-PH) method was used to analyze sPTB in the time-to-event analysis framework.8 The follow-up was from the time of measurement at around 28 wkGA until sPTB (the  \nevent) or censoring. All iPTBs were censored at the time of birth and term births were censored at 37 wkGA. The weighted Cox-PH was used to investigate the effect of several metabolites upon the time to sPTB while accounting for the case-cohort design. The design was specified with the twophase function in the survey R package.9 The svycoxph function then fitted a PH Cox model with Lin-Ying weights, accounting for the design.10,11 Extensions of correlation tests of Schoenfeld residuals and event time to a case-cohort setting were implemented to test the PH assumption.12 A cross-validation based method was used to estimate the survival (i.e. continuation of pregnancy) distribution of risk groups.13 The relationship between metabolites and sPTB was then examined by dividing samples into four risk groups by risk score quartile. The time-to-event model was evaluated using the C-index. To test the null hypothesis of no d","cbCaijegi5a3TW4l","https://ap.wps.com/l/cbCaijegi5a3TW4l","pdf",1627044,1,16,"English","en",105,"# Supplementary Methods\n## Supplementary Methods - Tables and Figures\n## Supplementary References\n## Software\n## Classification models\n## Prognosis prediction\n## Feature selection\n## Internal validation","[{\"question\":\"How were metabolite values prepared before model training?\",\"answer\":\"Metabolite values (MoMs) were log-transformed to address skewness, and metabolites with zero or near-zero variance in the study population were excluded before analysis.\"},{\"question\":\"Which supervised machine learning algorithms were used to classify sPTB samples?\",\"answer\":\"Six algorithms were applied: GBM, LDA, penalized logistic regression (LOG), random forest (RF), recursive partitioning and regression analysis (Rpart), and support vector machine (SVM).\"},{\"question\":\"How was internal validation of prediction performance performed?\",\"answer\":\"Models were validated using penalized logistic regression with an elastic net penalty, best-subsets selection based on AIC/BIC, and optimism-corrected AUC computed via 10-fold cross-validation with 100 repeats.\"}]","Prediction of Spontaneous Preterm Birth Using Supervised Machine Learning on Metabolomic Data - Supplementary Methods | PDF",1785894082,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prediction-of-spontaneous-preterm-birth-using-supervised-machine-learning-on-metabolomic-data-supplementary-methods","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/prediction-of-spontaneous-preterm-birth-using-supervised-machine-learning-on-metabolomic-data-supplementary-methods/124717/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How were metabolite values prepared before model training?","Question",{"text":75,"@type":76},"Metabolite values (MoMs) were log-transformed to address skewness, and metabolites with zero or near-zero variance in the study population were excluded before analysis.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which supervised machine learning algorithms were used to classify sPTB samples?",{"text":80,"@type":76},"Six algorithms were applied: GBM, LDA, penalized logistic regression (LOG), random forest (RF), recursive partitioning and regression analysis (Rpart), and support vector machine (SVM).",{"name":82,"@type":73,"acceptedAnswer":83},"How was internal validation of prediction performance performed?",{"text":84,"@type":76},"Models were validated using penalized logistic regression with an elastic net penalty, best-subsets selection based on AIC/BIC, and optimism-corrected AUC computed via 10-fold cross-validation with 100 repeats.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]