[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84724-en":3,"doc-seo-84724-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84724,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","WPG MoE Weak-Prior-Guided Dense Mixture-of-Experts for User-Level Social Media Depression Detection","Online social media posts offer scalable signals for early depression screening, yet common screening-stage designs often defer final decisions to a single detector. This overlooks heterogeneous user expression after screening: a monolithic classifier averages across users, diluting localized evidence and increasing errors, especially for non-self-disclosing users. WPG-MoE introduces weak-prior-guided dense mixture-of-experts on a shared LLM backbone. It learns routing using LUPI with LLM-extracted structured evidence for training-time priors, while inference relies on deployable PHQ-9 template screening plus the backbone. Experiments on Chinese and English datasets show improved performance with interpretable routing behavior.","WPG-MoE: Weak-Prior-Guided Dense Mixture-of-Experts for User-Level  \nSocial Media Depression Detection  \nXian Li 1 ,2 Yuanhe Tian2 * Yang Yang 1 Guoqing Wang 1 Yan Song3  \n1University of Electronic Science and Technology of China  \n2Zhongguancun Academy  \n3University of Science and Technology of China  \n[xianli@stu.uestc.edu.cn](xianli@stu.uestc.edu.cn) [yhtian94@gmail.com](yhtian94@gmail.com)  \n[yang.yang@uestc.edu.cn](yang.yang@uestc.edu.cn) [gqwang0420@uestc.edu.cn](gqwang0420@uestc.edu.cn) [clksong@gmail.com](clksong@gmail.com)  \narXiv :2607 .04350v 1 [ cs .CL] 5 Jul 2026  \nAbstract  \nOnline social media posts provide scalable signals for early depression screening, and recent studies mainly improve pre-classification evidence through risk-post selection, symptom grounding, and clinically informed feature construction. However, these screening-stage designs often leave final decisions to a single detector, overlooking how users heterogeneously express depressive risk after screening. A monolithic classifier must average across heterogeneous users, which may dilute localized evidence and cause misclassification, especially for non-self-disclosing users. To address this issue, we propose WPG-MoE, a weak-priorguided dense mixture-of-experts framework built on a shared large language model (LLM) backbone. WPG-MoE derives user-level weak semantic priors to softly route users to experts matched to different evidence layouts. We formulate this process as learning using privileged information (LUPI): rich LLM-extracted structured evidence guides training-time routing, while inference retains only Patient Health Questionnaire-9 (PHQ-9) template screening and the deployable backbone. Experiments on Chinese and English datasets show that WPGMoE outperforms strong baselines with interpretable routing behavior.  \n1 Introduction  \nDepression affects an estimated 332 million people worldwide, yet treatment coverage and minimally adequate care remain limited (Moitra et al., 2022 ; World Health Organization, 2025) . Social media histories therefore provide scalable early-identification signals, as De Choudhury et al.(2013) showed for depression detection from naturalistic user traces (Yates et al., 2017 ; Shen et al.,  \n* Corresponding Author.  \nFigure 1: Heterogeneous evidence patterns create a mismatch for monolithic modeling.  \n2017 ; Guntuku et al., 2017 ; Song et al., 2018 ; Chancellor and De Choudhury, 2020 ; Nie et al., 2020 ; Diao et al., 2023 ; Hu et al., 2024) .  \nRecent work improves user-level social media depression detection with complete-history modeling and clinically structured evidence: multimodal fusion, user-post summarization, symptomaware temporal modeling, capsule-style aggregation (Shen et al., 2017 ; Gui et al., 2019 ; Zoganet al., 2021 ; Wang et al., 2022 ; Cai et al., 2023 ; Liu et al., 2024), Patient Health Questionnaire-9 (PHQ-9) or psychiatric-scale guidance (Nguyen et al., 2022 ; Zhang et al., 2022 ; Wang et al., 2025), and large language model (LLM)-based annotation, summarization, retrieval, or explanation for clinical evidence (Wang et al., 2024 ; Lan et al., 2025 ;  \nTian et al., 2024a ; Ravenda et al., 2025) . Yet final predictors across pretrained language models (PLMs), sentence-embedding pipelines, capsule models, tree classifiers, and retrieval-augmented LLM agents still make one decision after screening.  \nThe bottleneck is post-screening heterogeneity, not evidence retrieval. Depressed users reveal risk through overlapping structures: diagnoses or medication, sustained symptoms, or a few high-intensity posts amid otherwise irrelevant histories (Mendes and Caseli, 2024) . These are evidence structures, not hard clinical subtypes. Figure 1 illustrates themismatch: after screening, a flat detector compresses heterogeneous signals into one representation and boundary. This averaging can dilute localized evidence and obscure weaker non-selfdisclosure patterns, motivating dense conditional specialization","cbCairytTkchQ91n","https://ap.wps.com/l/cbCairytTkchQ91n","pdf",2596645,2,1,23,"English","en",105,"# Abstract\n# 1 Introduction\n# 2 The Approach\n## 2.1 Problem Definition","[{\"question\":\"What limitation in current screening-stage designs does WPG-MoE address?\",\"answer\":\"It addresses the mismatch caused by heterogeneous post-screening evidence: final decisions are made by a single detector that averages across users, diluting localized evidence and harming accuracy for non-self-disclosing users.\"},{\"question\":\"How does WPG-MoE route users to experts?\",\"answer\":\"WPG-MoE softly routes users to experts based on user-level weak semantic priors that match different evidence layouts, using a shared LLM backbone.\"},{\"question\":\"What role does LUPI (learning using privileged information) play in training and inference?\",\"answer\":\"Training uses privileged, LLM-extracted structured evidence to build routing priors and evidence blocks, while inference keeps only deployable PHQ-9 template screening together with the shared backbone and history-level signals.\"}]",1784197865,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wpg-moe-weak-prior-guided-dense-mixture-of-experts-for-user-level-social-media-depression-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/wpg-moe-weak-prior-guided-dense-mixture-of-experts-for-user-level-social-media-depression-detection/84724/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-19","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation in current screening-stage designs does WPG-MoE address?","Question",{"text":75,"@type":76},"It addresses the mismatch caused by heterogeneous post-screening evidence: final decisions are made by a single detector that averages across users, diluting localized evidence and harming accuracy for non-self-disclosing users.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does WPG-MoE route users to experts?",{"text":80,"@type":76},"WPG-MoE softly routes users to experts based on user-level weak semantic priors that match different evidence layouts, using a shared LLM backbone.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does LUPI (learning using privileged information) play in training and inference?",{"text":84,"@type":76},"Training uses privileged, LLM-extracted structured evidence to build routing priors and evidence blocks, while inference keeps only deployable PHQ-9 template screening together with the shared backbone and history-level signals.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]