[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128288-en":3,"doc-seo-128288-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128288,962085570644,"Evangeline","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Discovering research articles containing evolutionary timetrees by machine learning","The study addresses the growing challenge of automatically identifying scientific articles that include evolutionary timetrees for inclusion in the TimeTree of Life resource. A text-mining and fine-tuned BERT approach was developed to classify full-text articles and improve detection by leveraging excerpts around mentioned figures. The optimized system is implemented in TimeTreeFinder (TTF) to process millions of papers, targeting higher F1 performance and increased coverage, while reducing manual cost and time.","Bioinformatics, 39(1), 2023, btad035  \n[https://doi.org/10.1093/bioinformatics/btad035](https://doi.org/10.1093/bioinformatics/btad035)[ ](https://doi.org/10.1093/bioinformatics/btad035)Advance Access Publication Date: 17 January 2023 Original Paper  \nData and text mining  \nDiscovering research articles containing evolutionary timetrees by machine learning  \nMarija Stanojevic  1, *, Jovan Andjelkovic 1, Adrienne Kasprowicz2, Louise A. Huuki2, Jennifer Chao 1, S. Blair Hedges2,3, Sudhir Kumar  2,3, * and Zoran Obradovic 1, *  \n1Center for Data Analytics and Biomedical Informatics, Computer and Information Science Department, Temple University, Philadelphia, PA 19121, USA, 2Department of Biology, Temple University, Philadelphia, PA 19121, USA and 3 Institute for Genomics and Evolutionary Medicine, Temple University, Philadelphia, PA 19121, USA  \n*To whom correspondence should be addressed.  \nAssociate Editor: Jonathan Wren  \nReceived on October 30, 2022; revised on January 7, 2023; editorial decision on January 12, 2023; accepted on January 16, 2023  \nAbstract  \nMotivation: Timetrees depict evolutionary relationships between species and the geological times of their divergence. Hundreds of research articles containing timetrees are published in scientiﬁc journals every year. The TimeTree (TT) project has been manually locating, curating and synthesizing timetrees from these articles for almost two decades into a TimeTree of Life, delivered through a unique, user-friendly web interface ([timetree.org](timetree.org)). The manual process of ﬁnding articles containing timetrees is becoming increasingly expensive and time-consuming. So, we have explored the effectiveness of text-mining approaches and developed optimizations to ﬁnd research articles containing timetrees automatically.  \nResults: We have developed an optimized machine learning system to determine if a research article contains an evolutionary timetree appropriate for inclusion in the TT resource. We found that BERT classiﬁcation ﬁne-tuned on whole-text articles achieved an F1 score of 0.67, which we increased to 0.88 by text-mining article excerpts surrounding the mentioning of ﬁgures. The new method is implemented in the TimeTreeFinder (TTF) tool, which automatically processes millions of articles to discover timetree-containing articles. We estimate that the TTF tool would produce twice as many timetree-containing articles as those discovered manually, whose inclusion in the TT database would potentially double the knowledge accessible to a wider community. Manual inspection showed that the precision on out-of-distribution recently published articles is 87% . This automation will speed up the collection and curation of timetrees with much lower human and time costs.  \nAvailability and implementation: [https://github.com/marija-stanojevic/time-tree-classi](https://github.com/marija-stanojevic/time-tree-classi)ﬁcation.  \n[Contact:](Contact: marija.stanojevic@temple.edu)[ marija.stanojevic@temple.edu](Contact: marija.stanojevic@temple.edu) [or s.kumar@temple.edu](or s.kumar@temple.edu) or [zoran.obradovic@temple.edu](zoran.obradovic@temple.edu)  \n[Supplementary information:](Supplementary information: Supplementary data are available at Bioinformatics online.)[ Supplementary data](Supplementary information: Supplementary data are available at Bioinformatics online.)[ are available at](Supplementary information: Supplementary data are available at Bioinformatics online.)[ Bioinformatics](Supplementary information: Supplementary data are available at Bioinformatics online.)[ online.](Supplementary information: Supplementary data are available at Bioinformatics online.)  \n1 Introduction  \nText mining has the potential to reshape the discovery of research articles containing domain-specific knowledge from the fastexpanding corpus of scientific literature. One such area is evolutionary biology, in which the growing affordability of genome sequencing technology has revolution","cbCaipZftVHdGK65","https://ap.wps.com/l/cbCaipZftVHdGK65","pdf",6999144,1,7,"English","en",105,"# Introduction\n## Motivation and background\n# Methods\n## Machine learning system design\n# Results\n## Classification performance and improvements\n## Impact on TimeTree coverage","[{\"question\":\"How much does the approach improve detection performance and coverage?\",\"answer\":\"Whole-text BERT classification achieves an F1 score of 0.67, which increases to 0.88 by focusing on figure-referencing excerpts. The authors estimate TTF could produce about twice as many timetree-containing articles as manual discovery, potentially doubling knowledge accessible via the TimeTree database.\"}]","Discovering research articles containing evolutionary timetrees by machine learning | PDF",1785946579,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"discovering-research-articles-containing-evolutionary-timetrees-by-machine-learning","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/discovering-research-articles-containing-evolutionary-timetrees-by-machine-learning/128288/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How much does the approach improve detection performance and coverage?","Question",{"text":75,"@type":76},"Whole-text BERT classification achieves an F1 score of 0.67, which increases to 0.88 by focusing on figure-referencing excerpts. The authors estimate TTF could produce about twice as many timetree-containing articles as manual discovery, potentially doubling knowledge accessible via the TimeTree database.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]