[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121679-en":3,"doc-seo-121679-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},121679,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Improved prediction of drug-induced liver injury literature using natural language processing and machine learning methods","Drug-induced liver injury (DILI) is a serious adverse hepatic drug reaction that may progress to life-threatening liver failure. Manual curation of DILI-related publications in PubMed is labor-intensive and can limit recall. To automate literature identification, an integrated natural language processing and machine learning classification model was developed using only paper titles and abstracts. A linear SVM trained on CAMDA challenge data combined TF-IDF and Word2Vec features, achieving strong internal and external test performance.","TYPE Original Research PUBLISHED 17 July 2023  \nDOI 10.3389/fgene.2023.1161047  \nOPEN ACCESS  \nEDITED BY  \nJoaquin Dopazo,  \nJunta de Andalucía, Spain  \nREVIEWED BY  \nRunzhi Zhang, AbbVie, United States Yogesh Sabnis,  \nUCB Biopharma SPRL, Belgium Minjun Chen,  \nNatioSnal Center for Toxicological Research (FDA), United States  \n*CORRESPONDENCE  \nJung Hun Oh,  [ohj@mskcc.org](ohj@mskcc.org)  \nRECEIVED 01 March 2023  \nACCEPTED 29 June 2023  \nPUBLISHED 17 July 2023  \nCITATION  \nOh JH, Tannenbaum A and Deasy JO (2023), Improved prediction of druginduced liver injury literature using natural language processing and machine learning methods.  \nFront. Genet. 14:1161047 .  \ndoi: 10.3389/fgene.2023.1161047  \nCOPYRIGHT  \n© 2023 Oh, Tannenbaum and Deasy. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nImproved prediction of  \ndrug-induced liver injury literature using natural language processing and machine learning methods  \nJung Hun Oh 1*, Allen Tannenbaum 2,3 and Joseph O. Deasy 1  \n1Department of Medical Physics, Memorial Sloan Kettering Cancer Center, New York, NY, United States, 2Department of Computer Science, Stony Brook University, Stony Brook, NY, United States, 3Department of Applied Mathematics and Statistics, Stony Brook University, Stony Brook, NY, United States  \nDrug-induced liver injury (DILI) is an adverse hepatic drug reaction that can potentially lead to life-threatening liver failure. Previously published work in thescientiﬁc literature on DILI has provided valuable insights for the understanding of hepatotoxicity as well as drug development. However, the manual search ofscientiﬁc literature in PubMed is laborious and time-consuming. Natural language processing (NLP) techniques along with artiﬁcial intelligence/machine learning approaches may allow for automatic processing in identifying DILIrelated literature, but useful methods are yet to be demonstrated. To address this issue, we have developed an integrated NLP/machine learning classiﬁcation model to identify DILI-related literature using only paper titles and abstracts. For prediction modeling, we used 14,203 publications provided by the Critical Assessment of Massive Data Analysis (CAMDA) challenge, employing word vectorization techniques in NLP in conjunction with machine learning methods. Classiﬁcation modeling was performed using 2/3 of the data for training and the remainder for test in internal validation. The best performance was achieved using a linear support vector machine (SVM) model on the combined vectors derived from term frequency-inverse document frequency (TF-IDF) and Word2Vec, resulting in an accuracy of 95 . 0% and an F1-score of 95 . 0% . The ﬁnal SVM model constructed from all 14,203 publications was tested on independent datasets, resulting in accuracies of 92.5%, 96.3%, and 98.3%, and F1-scores of 93. 5%, 86 . 1%, and 75 . 6% for three test sets (T1-T3) . Furthermore, the SVM model was tested on four external validation sets (V1-V4), resulting in accuracies of 92. 0%, 96 . 2%, 98 .3%, and 93 . 1%, and F1-scores of 92 .4%, 82 . 9%, 75 . 0%, and 93 .3% .  \nKEYWORDS  \nnatural language processing, TF-IDF, Word2vec, drug-induced liver injury, artiﬁcial intelligence, machine learning  \n1 Introduction  \nDrug-induced liver injury (DILI) is a liver disease caused by an adverse drug reaction that can potentially lead to fatal liver failure (David and Hamilton, 2010) . Previously published work in the scientiﬁc literature on DILI has provided valuable insights for the understanding of hepatotoxicity on causative agents and clinical features as well as dr","cbCaie8jF2MYec7A","https://ap.wps.com/l/cbCaie8jF2MYec7A","pdf",2084129,1,"English","en",105,"# Introduction\n## Background on DILI and the need for literature retrieval\n## NLP and word vectorization for feature extraction\n## Word embedding methods and text classification\n## Transformer-based language models (BERT)","[{\"question\":\"Why is automated prediction of DILI-related literature needed?\",\"answer\":\"Manual PubMed searches for DILI are time-consuming and can reduce recall, limiting the ability to comprehensively identify relevant hepatotoxicity publications.\"},{\"question\":\"What data and inputs does the proposed model use?\",\"answer\":\"The model identifies DILI-related literature using only paper titles and abstracts as input features.\"},{\"question\":\"Which modeling approach and features produced the best results?\",\"answer\":\"A linear support vector machine (SVM) using combined TF-IDF and Word2Vec representations achieved the highest performance, reaching about 95% accuracy and F1-score in internal evaluation and strong results on external validation sets.\"}]","Improved prediction of drug-induced liver injury literature using natural language processing and machine learning methods | PDF",1785806169,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"improved-prediction-of-drug-induced-liver-injury-literature-using-natural-language-processing-and-machine-learning-methods","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/improved-prediction-of-drug-induced-liver-injury-literature-using-natural-language-processing-and-machine-learning-methods/121679/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is automated prediction of DILI-related literature needed?","Question",{"text":74,"@type":75},"Manual PubMed searches for DILI are time-consuming and can reduce recall, limiting the ability to comprehensively identify relevant hepatotoxicity publications.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What data and inputs does the proposed model use?",{"text":79,"@type":75},"The model identifies DILI-related literature using only paper titles and abstracts as input features.",{"name":81,"@type":72,"acceptedAnswer":82},"Which modeling approach and features produced the best results?",{"text":83,"@type":75},"A linear support vector machine (SVM) using combined TF-IDF and Word2Vec representations achieved the highest performance, reaching about 95% accuracy and F1-score in internal evaluation and strong results on external validation sets.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]