[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125580-en":3,"doc-seo-125580-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125580,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Identifying potentially excellent publications using a citation-based machine learning approach - Paper","Excellent research publications drive advances in science and technology, yet the early discovery of potentially excellent papers remains challenging due to the rapid growth of articles. This study develops citation-based machine learning to identify potentially excellent papers (PEPs) in artificial intelligence, using 5 static and 8 time-dependent citation features. It models Random Forest, LightGBM, Naive Bayes, SVM, neural networks, and TabNet, training and testing on 96,169 AI papers collected from Web of Science between 1990–2010. Results show time-dependent and citation-peak features matter most, threshold choice affects outcomes, and LightGBM performs best for accuracy and recall.","This is a repository copy of Identifying potentially excellent publications using a citationbased machine learning approach.  \nWhite Rose Research Online URL for this paper:  \n[https://eprints.whiterose.ac.uk/197983/](https://eprints.whiterose.ac.uk/197983/)  \nVersion: Published Version  \nArticle:  \nHu, Z. , Cui, J. and Lin, A. (2023) Identifying potentially excellent publications using a  \ncitation-based machine learning approach. Information Processing & Management, 60 (3) .  \n103323. ISSN 0306-4573  \n[https://doi.org/10.1016/j.ipm.2023.103323](https://doi.org/10.1016/j.ipm.2023.103323)  \nReuse  \nThis article is distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivs (CC BY-NC-ND) licence. This licence only allows you to download this work and share it with others as long as you credit the authors, but you can’t change the article in any way or use it commercially. More information and the full terms of the licence here: [https://creativecommons.org/licenses/](https://creativecommons.org/licenses/)  \nTakedown  \nIf you consider content in White Rose Research Online to be in breach of UK law, please notify us by  \nemailing [eprints@whiterose.ac.uk](eprints@whiterose.ac.uk) including the URL of the record and the reason for the withdrawal request.  \n[eprints@whiterose.ac.uk](eprints@whiterose.ac.uk)[ ](eprints@whiterose.ac.uk)[https://eprints.whiterose.ac.uk/](https://eprints.whiterose.ac.uk/)  \nInformation Processing and Management 60 (2023) 103323  \nContents lists available at ScienceDirect  \nInformation Processing and Management  \njournal [homepage:](homepage: www.elsevier.com/locate/infoproman)[ www.elsevier.com/locate/infoproman](homepage: www.elsevier.com/locate/infoproman)  \n| Identifying potentially excellent publications using a citation-based machine learning approach |  |  |  |\n| --- | --- | --- | --- |\n| Zewen Hu a, Jingjing Cui a, Angela Lin b, *\u003Cbr>a School of Management Science and Engineering, Nanjing University of Information Science and Technology, Nanjing 210044, China b Information School, University of Sheffield, Sheffield S10 2TN, United Kingdom |  |  |  |\n| A R T I C L E I N F O |  | A B S T R A C T |  |\n| Keywords:\u003Cbr>Machine learning Artificial intelligence Excellent papers Highly cited papers Sleeping beauty Citation-based measures\u003Cbr>Citation peak Neural network LightGBM TabNet |  | Excellent research papers are vital to science and technology advances. Thus, the early identification of potentially excellent research papers and recognizing their value in science and technology is high on the research agenda. This study used a set of 5 static and 8 time-dependent citation features to explore six machine learning methods and identify the method with the best performance to identify potentially excellent papers. The study modelled Random Forest, LightGBM, Naive Bayes, Support Vector Machine, Neural Network, and TabNet to identify PEPs in the artificial intelligence field. The study defined highly cited papers using the threshold of the top 1% and top 5% and collected the data from the Web of Science®. Bibliometric and citation data from 485,041 research articles, proceeding papers, and reviews published in AI between 1990 and 2010 were collected initially. The data was screened and processed, and the final dataset consists of 96,169 papers for the training and test sets. The findings suggest that the timedependent citation features are more important than the static features, and citation peak features are more significant than the citation features in identifying potentially excellent papers. The findings demonstrate the effect of threshold on machine learning outcomes (e.g., the top 1% and 5%); therefore, the study argues that the decision about threshold selection should be carefully made. LightGBM and Random Forest both performed with the given conditions and achieved the same score in accuracy and recall. Nevertheless, when comparing their performance in other indicato","cbCaisB4aC0CidOg","https://ap.wps.com/l/cbCaisB4aC0CidOg","pdf",1408102,1,23,"English","en",105,"# Introduction\n## Problem and motivation\n## Definition challenges of “excellent” articles\n# Approach overview\n## Citation features (static and time-dependent)\n## Machine learning models evaluated\n# Data and experimental setup\n## Dataset collection and filtering\n## Training/test construction\n# Results and discussion\n## Feature importance findings\n## Threshold effects on performance\n## Model comparison and best performer","[{\"question\":\"What problem does the study address?\",\"answer\":\"It targets the early identification of potentially excellent research publications when the volume of papers makes discovery difficult.\"},{\"question\":\"How does the study define and detect “potentially excellent papers” (PEPs)?\",\"answer\":\"It uses citation-based measures, defining highly cited papers via thresholds (top 1% and top 5%) and then learning patterns to identify PEPs.\"},{\"question\":\"Which machine learning model performs best and why?\",\"answer\":\"LightGBM is concluded as the best-performing model, with strong accuracy and recall under the studied conditions, and better performance on additional indicators like F1 and cross-entropy loss.\"}]","Identifying potentially excellent publications using a citation-based machine learning approach - Paper | PDF",1785900003,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"identifying-potentially-excellent-publications-using-a-citation-based-machine-learning-approach-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/identifying-potentially-excellent-publications-using-a-citation-based-machine-learning-approach-paper/125580/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address?","Question",{"text":75,"@type":76},"It targets the early identification of potentially excellent research publications when the volume of papers makes discovery difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study define and detect “potentially excellent papers” (PEPs)?",{"text":80,"@type":76},"It uses citation-based measures, defining highly cited papers via thresholds (top 1% and top 5%) and then learning patterns to identify PEPs.",{"name":82,"@type":73,"acceptedAnswer":83},"Which machine learning model performs best and why?",{"text":84,"@type":76},"LightGBM is concluded as the best-performing model, with strong accuracy and recall under the studied conditions, and better performance on additional indicators like F1 and cross-entropy loss.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]