[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127199-en":3,"doc-seo-127199-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127199,549768702563,"Sage","https://ap-avatar.wpscdn.com/avatar/8000c4aa63b76e948b?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786536092046926083",8,"Research & Report","Predicting Clinical Trial Completion and Success Using Machine Learning and Natural Language Processing - May 2025","This study develops a dual-task machine learning framework to predict both the operational completion status and the scientific success of clinical trials using data from ClinicalTrials.gov. Models combine structured trial metadata with unstructured textual descriptions to estimate completion likelihood and whether trials meet primary endpoints. Ensemble methods such as XGBoost outperform standard baselines, especially with contextual embeddings from BioLinkBERT. For success labeling, the work introduces an LLM-driven GPT-4o-mini annotation pipeline, validated via human evaluation. Results show that integrating NLP and scalable LLM labeling improves forecasting and better leverages biomedical data.","THE UNIVERSITY OF CHICAGO  \nPredicting Clinical Trial Completion and Success Using Machine Learning and Natural Language  \nProcessing  \nBy  \nJiazheng Li  \nMay 2025  \nA paper submitted in partial fulfillment of the requirements for the Master of Arts degree in the Master of Arts in Computational Social Science  \nFaculty Advisor: Professor Yuan Ji  \nPreceptor: Fabricio Vasselai  \nAbstract  \nThis study introduces a dual-task machine learning framework for predicting both the opera[tional completion and scientific success of clinical trials using data from ClinicalTrials.gov. Leverag](tional completion and scientific success of clinical trials using data from ClinicalTrials.gov. Leverag)ing structured trial metadata and unstructured textual descriptions, we develop predictive models that assess whether trials are likely to complete and whether they meet their primary endpoints. For the first task, ensemble models like XGBoost significantly outperform traditional baselines, particularly when enriched with contextual embeddings derived from BioLinkBERT. For the second task, we propose a novel large language model (LLM)-driven annotation pipeline using GPT-4o-mini to label trial success based on publication content. Human evaluation confirms its high accuracy. Across both tasks, our framework demonstrates the value of combining structured features, natural language processing, and scalable LLM-based labeling to improve the understanding and forecasting of clinical trial performance. This approach not only enhances predictive accuracy but also contributes to better utilization of large-scale biomedical data.  \nKeywords: Clinical Trials; Machine Learning; Natural Language Processing; Clinical Trial Completion; Outcome Prediction; GPT-4; BioLinkBERT; [ClinicalTrials.gov](ClinicalTrials.gov) ; Publication Analysis  \n1 Introduction  \nClinical trials play a pivotal role in advancing medical knowledge by evaluating the safety and efficacy of new treatments, interventions, and diagnostic methods. As these trials are essential for validating medical innovations, predicting their outcomes can offer significant insights for researchers, practitioners, and policymakers. Despite the importance of clinical trials, many face challenges such as prolonged time frames, recruitment difficulties, and rising costs, often leading to premature termination or inconclusive results.  \nThis study seeks to predict two key outcomes: (1) the completion or termination status of clinical trials and (2) the effectiveness of trials, measured by whether the associated publication reports positive or negative results—as determined by a large language model (LLM) analyzing the publication abstract for evidence of primary endpoint achievement. By employing large language models (LLMs) to tokenize and analyze the associated published research, this study aims to uncover patterns and predictors that influence both trial completion and effectiveness.  \nThe primary data source for this research is the [ClinicalTrials.gov](ClinicalTrials.gov) database, a comprehensive repository that provides detailed information on clinical studies conducted around the world. Established under the Food and Drug Administration Amendments Act (FDAAA) of 2007, this database contains records of trials that meet specific criteria, such as those involving FDA-regulated drug or device products, studies with U.S. sites, or trials using products exported from the U.S. for research. The dataset includes key details about each trial, including its registration information, current status, sponsor, intervention type, and, when available, outcomes. Additionally, it tracks whether a study is completed, terminated prematurely, or still ongoing. With over half of millions of entries, ClinicalTrials.gov offers valuable data for predictive modeling, providing insights into the success, failure, and potential impact of clinical trials across a wide range of medical disciplines.  \nThis research builds on existi","cbCaiuTfGTP4tW6e","https://ap.wps.com/l/cbCaiuTfGTP4tW6e","pdf",1153314,1,40,"English","en",105,"# Introduction\n## Research motivation and target outcomes\n## Data source: ClinicalTrials.gov\n## Background on trial complexity and early signals\n# Literature Review\n## General landscape of clinical trials\n## Trial termination","[{\"question\":\"What two outcomes does the study aim to predict?\",\"answer\":\"The study predicts (1) whether a clinical trial is completed or terminated and (2) whether the trial is effective, assessed by whether its publication reports evidence of primary endpoint achievement using an LLM.\"},{\"question\":\"How does the framework use data for prediction?\",\"answer\":\"It combines structured trial metadata with unstructured textual descriptions from ClinicalTrials.gov, then trains models to assess completion status and endpoint-related scientific success.\"},{\"question\":\"How are trial success labels created for the effectiveness task?\",\"answer\":\"The study proposes an LLM-driven annotation pipeline using GPT-4o-mini to label trial success based on publication content, followed by human evaluation to confirm high accuracy.\"}]","Predicting Clinical Trial Completion and Success Using Machine Learning and Natural Language Processing - May 2025 | PDF",1785937455,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predicting-clinical-trial-completion-and-success-using-machine-learning-and-natural-language-processing-may-2025","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/predicting-clinical-trial-completion-and-success-using-machine-learning-and-natural-language-processing-may-2025/127199/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What two outcomes does the study aim to predict?","Question",{"text":75,"@type":76},"The study predicts (1) whether a clinical trial is completed or terminated and (2) whether the trial is effective, assessed by whether its publication reports evidence of primary endpoint achievement using an LLM.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the framework use data for prediction?",{"text":80,"@type":76},"It combines structured trial metadata with unstructured textual descriptions from ClinicalTrials.gov, then trains models to assess completion status and endpoint-related scientific success.",{"name":82,"@type":73,"acceptedAnswer":83},"How are trial success labels created for the effectiveness task?",{"text":84,"@type":76},"The study proposes an LLM-driven annotation pipeline using GPT-4o-mini to label trial success based on publication content, followed by human evaluation to confirm high accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":21,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]