[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84450-en":3,"doc-seo-84450-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84450,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782109480056885918",8,"Research & Report","Identifying Public Response Topics to CDC COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing","Public health agencies increasingly use social media to monitor responses to health communication during crises such as COVID-19, yet analyzing large-scale short replies is difficult because the text is brief, informal, and highly variable. Topic modeling can reveal thematic patterns, but its results are often low quality on noisy social media data. This study introduces TM-Rephrase, a model-agnostic framework that uses large language model–based rephrasing to improve topic quality, interpretability, and semantic relevance for CDC communications on X.","Original Paper  \nIdentifying Public Response Topics to CDC COVID-19 Communications on Social Media: Infoveillance Study Using Large Language Model–Based Rephrasing  \nWangjiaxuan Xin 1 *, Shuhua Yin2, Shi Chen2, Yaorong Ge 1  \n1College of Computing and Informatics, University of North Carolina at Charlotte, Charlotte, North Carolina, NC, United States  \n2Department of Epidemiology and Community Health, University of North Carolina at Charlotte, Charlotte, NC, United States  \n*Corresponding Author: Wangjiaxuan Xin  \nUniversity of North Carolina at Charlotte 9201 University City Blvd Charlotte, NC, 28223  \nUnited States  \nPhone: +1 7048776536  \n[Email: wxin@charlotte.edu](Email: wxin@charlotte.edu)  \nAbstract  \nBackground: Public health agencies increasingly rely on social media platforms such as X (formerly Twitter) to monitor public responses to health communications and to support timely, data-driven decision-making during health crises such as the COVID-19 pandemic. In particular, public replies to official communications from agencies such as the Centers for Disease Control and Prevention (CDC) provide valuable insights into population-level perceptions, concerns, and engagement with health policies. However, effectively analyzing these large-scale short-text responses remains challenging due to their brevity, informality, and linguistic variability. Topic modeling offers a scalable approach to identifying thematic patterns in such data, but its performance is often limited when applied to short, noisy social media text, resulting in low-quality and difficult-to-interpret topics. Although recent advances in large language models (LLMs) provide new opportunities to enhance text representation, their potential to improve topic modeling for public health–oriented social media analysis remains underexplored.  \nObjective: This study aims to enhance public health surveillance by improving the analysis of public responses to official health communications on social media. To achieve this, we develop and evaluate TM-Rephrase, a model-agnostic framework that leverages LLM–based rephrasing to improve the quality, interpretability, and semantic relevance of topics derived by topic models for public health-related short texts on social media.  \nMethods: We analyzed 25,027 public replies to official CDC posts on X collected between May 2020 and November 2022 to examine public responses to health communications. We applied a LLM–based rephrasing framework (TM-Rephrase) to transform informal short texts into more standardized and context-enriched representations using general and colloquial-to-formal schemes. Both original and rephrased texts were analyzed across multiple topic models, and topic quality  \nwas evaluated using coherence, uniqueness, redundancy, and diversity metrics, along with qualitative assessment of interpretability and post-topic semantic alignment.  \nResults: TM-Rephrase consistently improved topic quality across models and evaluation metrics. For LDA, topic coherence increased from 􀜥􀯩 =0.3094 (no rephrasing) to 􀜥􀯩 =0.5004 with colloquial-to-formal rephrasing (>60% relative improvement) . For BERTopic, coherence improved from 􀜥􀯩 =0.4078 to 􀜥􀯩 =0.4734 . Diversity-related metrics also improved, with TSCTM achieving TU=1.0, TD=1.0, and TR=0 under rephrased conditions, indicating fully distinct topics. These improvements were robust across multiple LLMs (Gemini, GPT-4o-mini, and Mistral-7B) . Qualitative results further showed that rephrased texts produced more interpretable and semantically coherent topics (representative keywords), enabling clearer identification of public concerns and responses to health communications, including themes related to vaccination attitudes, perceived risks, and public health measures. This improvement facilitates more reliable characterization of public discourse, which is critical for supporting social media–based public health surveillance and informing communication strategies.  \nConclus","cbCaikkX3yuNtO9C","https://ap.wps.com/l/cbCaikkX3yuNtO9C","pdf",4760055,2,1,22,"English","en",105,"# Abstract\n## Background\n## Objective\n## Methods\n## Results\n## Conclusions\n# Introduction\n## Background","[{\"question\":\"What problem does the study address in analyzing CDC COVID-19 replies on social media?\",\"answer\":\"Short social media responses are difficult to analyze because they are brief, informal, and linguistically variable. This often limits the quality and interpretability of topic modeling results.\"},{\"question\":\"How does TM-Rephrase improve topic modeling for public health short texts?\",\"answer\":\"TM-Rephrase applies LLM-based rephrasing to transform informal short replies into more standardized, context-enriched representations. It then re-runs multiple topic models on both original and rephrased text to compare topic quality.\"},{\"question\":\"What improvements were observed in topic quality after rephrasing?\",\"answer\":\"Across models, TM-Rephrase improves coherence and diversity metrics, and qualitative evaluation shows topics become more interpretable with semantically coherent representative keywords. The paper links the improved topics to clearer identification of vaccination attitudes, perceived risks, and public health measures.\"}]",1784195696,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"identifying-public-response-topics-to-cdc-covid-19-communications-on-social-media-infoveillance-study-using-large-language-modelbased-rephrasing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/identifying-public-response-topics-to-cdc-covid-19-communications-on-social-media-infoveillance-study-using-large-language-modelbased-rephrasing/84450/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in analyzing CDC COVID-19 replies on social media?","Question",{"text":75,"@type":76},"Short social media responses are difficult to analyze because they are brief, informal, and linguistically variable. This often limits the quality and interpretability of topic modeling results.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TM-Rephrase improve topic modeling for public health short texts?",{"text":80,"@type":76},"TM-Rephrase applies LLM-based rephrasing to transform informal short replies into more standardized, context-enriched representations. It then re-runs multiple topic models on both original and rephrased text to compare topic quality.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements were observed in topic quality after rephrasing?",{"text":84,"@type":76},"Across models, TM-Rephrase improves coherence and diversity metrics, and qualitative evaluation shows topics become more interpretable with semantically coherent representative keywords. The paper links the improved topics to clearer identification of vaccination attitudes, perceived risks, and public health measures.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]