[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81485-en":3,"doc-seo-81485-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81485,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Topic Model Based on Co-occurrence Word Networks for Unbalanced Short Text Datasets","A straightforward solution is presented for detecting scarce topics in unbalanced short-text datasets. The proposed method, CWUTM, mitigates sparse and unbalanced topic behavior by reducing the influence of incidental word co-occurrence. CWUTM models topic distribution for each word through co-occurrence word networks, adjusts node activity computation, and partially normalizes representations across scarce and abundant topics to improve low-frequency topic sensitivity. Gibbs sampling, similar to LDA, supports flexible deployment. Experiments show superior scarce-topic discovery and effective early detection of emerging topics or unexpected events on social platforms.","Topic Model Based on Co-occurrence Word Networks for Unbalanced Short Text Datasets  \nChengjie Ma  \n[macj163@163.com](macj163@163.com)  \nBeijing Key Laboratory of Intelligent Telecommunication Software and Multimedia, Beijing University of Posts and Telecommunications  \nBeijing, China  \nJunping Du∗ [junpingdu@126.com](junpingdu@126.com)  \nBeijing Key Laboratory of Intelligent Telecommunication Software and Multimedia, Beijing University of Posts and Telecommunications  \nBeijing, China  \nMeiyu Liang  \n[meiyu1210@bupt.edu.cn](meiyu1210@bupt.edu.cn)  \nBeijing Key Laboratory of Intelligent Telecommunication Software and Multimedia, Beijing University of Posts and Telecommunications  \nBeijing, China  \nZeli Guan  \n[guanzeli@bupt.edu.cn](guanzeli@bupt.edu.cn)  \nBeijing Key Laboratory of Intelligent Telecommunication Software and Multimedia, Beijing University of Posts and Telecommunications  \nBeijing, China  \narXiv :2311 .02566v2 [ cs .CL] 10 Jul 2026  \nAbstract  \nWe propose a straightforward solution for detecting scarce topics in unbalanced short-text datasets. Our approach, named CWUTM (Topic model based on co-occurrence word networks for unbalanced short text datasets), addresses the challenge of sparse and unbalanced short text topics by mitigating the effects of incidental word co-occurrence. This allows our model to prioritize the identification of scarce topics (low-frequency topics) . Unlike previous methods, CWUTM leverages co-occurrence word networks to capture the topic distribution of each word, and enhances the sensitivity in identifying scarce topics by redefining the calculation of node activity and normalizing the representation of both scarce and abundant topics to some extent. Moreover, CWUTM adopts Gibbs sampling, similar to LDA, making it easily adaptable to various application scenarios. Extensive experimental validation on unbalanced short-text datasets demonstrates the superiority of CWUTM compared to baseline approaches in discovering scarce topics. According to the experimental results, the proposed model is effective in early and accurate detection of emerging topics or unexpected events on social platforms.  \nKeywords  \nScarce topic, co-occurrence network, unbalanced datasets  \n1 Introduction  \nTopic models serve as statistical tools used to uncover concealed semantic structures within document collections [1]. These models, along with their extensions, have found applications in diverse fields including marketing, sociology, political science, among others [2] . In applied intelligent decision scenarios, interpretable machine learning also emphasizes that learned semantic features should remain transparent and actionable [3] .  \n∗ Corresponding author.  \nThis work was supported by the Program of the National Natural Science Foundation of China (62192784, U22B2038, 62172056) and by Young Elite Scientists Sponsorship Program by CAST (2022QNRC001) .  \nMost topic models are advancements built upon the latent Dirichlet allocation (LDA) technique [1]. LDA, being the canonical form of existing topic models, employs a hierarchical parametric Bayesian approach to uncover topics within extensive corpora. It represents documents as a mixture of topics, with each topic being a probability distribution of words from the corpus vocabulary. Through statistical inference, LDA learns the probability distribution of words associated with each topic and the topic distribution for each document. However, LDA-like models, which leverage document-level word co-occurrence information [4], tend to consolidate semantically related words into a single topic. This characteristic makes them highly sensitive to the length and quantity of documents attributed to each topic. Consequently, when working with short texts that contain a limited number of words, these models fail to capture the relationships among words.  \nWith the rapid evolution of the World Wide Web and the emergence of various web applications, short texts have become p","cbCainTBZnl2adRi","https://ap.wps.com/l/cbCainTBZnl2adRi","pdf",479663,1,7,"English","en",105,"# Introduction\n## Topic modeling background and motivations\n## Limitations of LDA and co-occurrence-based variants\n## Short-text topic modeling challenges\n## Related approaches (pseudo-documents, external knowledge, modified LDA, neural methods, BITERM)","[{\"question\":\"What problem does CWUTM address in unbalanced short-text datasets?\",\"answer\":\"CWUTM focuses on detecting scarce topics in settings where topics are sparse and unevenly distributed, which makes low-frequency topics hard to learn and distinguish in short texts.\"},{\"question\":\"How does CWUTM differ from LDA-like topic models?\",\"answer\":\"CWUTM uses co-occurrence word networks to characterize the topic distribution of each word and redesigns node activity calculation, aiming to reduce the impact of incidental co-occurrence on scarce topic detection.\"},{\"question\":\"What inference method does CWUTM use and why is it useful?\",\"answer\":\"CWUTM adopts Gibbs sampling, similar to LDA, which makes the model easier to adapt to different application scenarios.\"}]",1784173761,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"topic-model-based-on-co-occurrence-word-networks-for-unbalanced-short-text-datasets","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/topic-model-based-on-co-occurrence-word-networks-for-unbalanced-short-text-datasets/81485/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CWUTM address in unbalanced short-text datasets?","Question",{"text":75,"@type":76},"CWUTM focuses on detecting scarce topics in settings where topics are sparse and unevenly distributed, which makes low-frequency topics hard to learn and distinguish in short texts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CWUTM differ from LDA-like topic models?",{"text":80,"@type":76},"CWUTM uses co-occurrence word networks to characterize the topic distribution of each word and redesigns node activity calculation, aiming to reduce the impact of incidental co-occurrence on scarce topic detection.",{"name":82,"@type":73,"acceptedAnswer":83},"What inference method does CWUTM use and why is it useful?",{"text":84,"@type":76},"CWUTM adopts Gibbs sampling, similar to LDA, which makes the model easier to adapt to different application scenarios.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]