[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84741-en":3,"doc-seo-84741-105":29,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84741,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Evaluation and Explainability of Unsupervised Scholarly Collaboration Recommendations","Evaluation of unsupervised scholarly collaboration recommendations using researchers’ publication text focuses on content-only signals without graph supervision. Three method families are compared: a TF-IDF lexical baseline, topic modeling with LDA and BERTopic (including clone variants), and embedding-based retrieval using SciBERT representations with Faiss. A constrained setting partially removes publication overlap while still using historical co-authorship as post-hoc ground truth. TF-IDF degrades sharply under reduced overlap, while topic and embedding approaches remain more stable, indicating reliance on broader distributional similarity. Explainability combines intrinsic topic interpretations and post-hoc retrieval-based rationales via language models, balancing transparency and readability.","Evaluation and Explainability of Unsupervised Scholarly Collaboration Recommendations  \n1st Md Asaduzzaman Noor Gianforte School of Computing Montana State University Bozeman, Montana, USA [mdasaduzzamannoor@montana.edu](mdasaduzzamannoor@montana.edu)  \n2nd John W. Sheppard Gianforte School of Computing Montana State University Bozeman, Montana, USA [john.sheppard@montana.edu](john.sheppard@montana.edu)  \n3rd Jason A. Clark Library  \nMontana State University Bozeman, Montana, USA [jaclark@montana.edu](jaclark@montana.edu)  \narXiv :2607 .04529v 1 [ cs .IR] 5 Jul 2026  \nAbstract—In this paper, we examine unsupervised, contentbased collaboration recommendations using publication text in scholarly settings. We compare three families of methods: a TFIDF baseline, topic-based models (LDA and BERTopic, including clone variants), and embedding-based retrieval using SciBERT with Faiss. To evaluate model behavior beyond simple lexical matching, we introduce a constrained setting where publication overlap between researchers is partially removed while still using historical co-authorship as proxy ground truth for post-hoc evaluation. Results show clear differences across methods. TFIDF performs best under full information but drops significantly as overlap is reduced. In contrast, topic-based and embeddingbased approaches show more stable performance, suggesting they capture broader distributional similarities, rather than relying only on direct lexical overlap. We also examine explainability through two perspectives: intrinsic topic-based explanations and post-hoc, retrieval-based explanations generated using language models. These provide complementary trade-offs between transparency and human readability.  \nIndex Terms—unsupervised recommendation, scholarly recommender systems, topic modeling, embedding-based retrieval  \nI. INTRODUCTION  \nA perennial challenge in scholarly settings is helping scholars identify potential research collaborators. In this work, we study unsupervised scholarly collaboration recommendation based solely on publication text (titles and abstracts) . The goal is to identify potential research collaborations using only content-based signals, without relying on prior graph supervision during model construction. Collaboration recommendation is particularly challenging in interdisciplinary settings, where historical co-authorship is sparse or absent, despite strong topical similarity.  \nWe evaluate three families of approaches. First, a TFIDF-based baseline is used as a lexical similarity method without any graph structure. Second, we use topic modelbased approaches (namely, LDA and BERTopic), along with their cloning variants, which are designed to better capture secondary research themes and mitigate publication imbalance across researchers. Third, we include an embedding-based retrieval method using SciBERT representations with Faiss for nearest-neighbor search over researcher-level embeddings.  \nEvaluation of unsupervised, content-based collaboration methods remains limited. While topic-based approaches provide interpretable recommendations, their evaluation often  \nfocuses on recovering observed collaborations, with less attention to how these methods behave under reduced or incomplete publication information.  \nA key aspect of our study is the evaluation setting. We use historical co-authorship links as a proxy ground truth for post-hoc evaluation, while ensuring that all models remain fully unsupervised during training. To better understand how models behave with limited information, we further evaluate performance under reduced publication overlap.  \nFor explainability, we consider two complementary approaches. Topic-based methods offer intrinsic interpretability through topic distributions, shared concepts, and word cloudbased visualizations. In contrast, for embedding-based methods, we introduce a post-hoc explanation framework that uses retrieved publication contexts and large language model-based summar","cbCaifdSmP5gDpTK","https://ap.wps.com/l/cbCaifdSmP5gDpTK","pdf",426371,1,6,"English","en",105,"# Introduction\n# Related Work\n# Evaluation Setup and Dataset Statistics\n# Explainability Approaches","[{\"question\":\"What problem does the paper address in scholarly settings?\",\"answer\":\"The paper targets identifying potential research collaborators using only publication text, especially when historical co-authorship is sparse or missing across disciplines.\"},{\"question\":\"Which recommendation methods are evaluated?\",\"answer\":\"It compares a TF-IDF lexical baseline, topic-model approaches (LDA and BERTopic with clone variants), and an embedding-based retrieval method using SciBERT with Faiss for nearest-neighbor search.\"},{\"question\":\"How does the paper evaluate performance beyond simple lexical matching?\",\"answer\":\"It introduces a constrained post-hoc evaluation where publication overlap between researchers is partially removed, while historical co-authorship links are used as proxy ground truth.\"},{\"question\":\"How is explainability handled for different method families?\",\"answer\":\"Topic-based methods provide intrinsic interpretability through topic distributions and related visualizations, while embedding-based methods use a post-hoc retrieval-and-language-model framework to generate human-readable rationales.\"}]",1784197992,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":27},"evaluation-and-explainability-of-unsupervised-scholarly-collaboration-recommendations","",{"@graph":35,"@context":89},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/evaluation-and-explainability-of-unsupervised-scholarly-collaboration-recommendations/84741/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in scholarly settings?","Question",{"text":75,"@type":76},"The paper targets identifying potential research collaborators using only publication text, especially when historical co-authorship is sparse or missing across disciplines.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which recommendation methods are evaluated?",{"text":80,"@type":76},"It compares a TF-IDF lexical baseline, topic-model approaches (LDA and BERTopic with clone variants), and an embedding-based retrieval method using SciBERT with Faiss for nearest-neighbor search.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper evaluate performance beyond simple lexical matching?",{"text":84,"@type":76},"It introduces a constrained post-hoc evaluation where publication overlap between researchers is partially removed, while historical co-authorship links are used as proxy ground truth.",{"name":86,"@type":73,"acceptedAnswer":87},"How is explainability handled for different method families?",{"text":88,"@type":76},"Topic-based methods provide intrinsic interpretability through topic distributions and related visualizations, while embedding-based methods use a post-hoc retrieval-and-language-model framework to generate human-readable rationales.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,118,123,126,131,134,138],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},"Technology",50,"technology",{"id":119,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":121,"slug":122},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":124,"slug":125},30,"research-report",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":129,"slug":130},9,"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":45,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]