[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86514-en":3,"doc-seo-86514-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86514,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization","Opinionated text across product reviews, hotel feedback, and social posts provides signals about user experiences, preferences, and concerns, but large, redundant, and imbalanced corpora hinder faithful opinion summarization. A proposed framework preserves semantics in LLM-based opinion summarization while minimizing token usage by combining multidimensional classification (sentiment, topics, emotions) with stratified sampling to select compact, representative, facet-balanced subsets. Tailored prompts then generate balanced, diversity-aware summaries. Experiments on Amazon, Tripadvisor, and X/Twitter show reduced cost and consistent gains in coverage, balance, and semantic preservation.","arXiv :2607 . 10825v 1 [ cs .CL] 12 Jul 2026  \nLarge Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization  \nFabrizio Marozzo University of Calabria, Rende, Italy  \n[fmarozzo@dimes. unical. it](fmarozzo@dimes. unical. it)  \nStefano Iannicelli University of Calabria, Rende, Italy  \nAbstract  \nOpinionated text—spanning product reviews, hotel feedback, and social posts—captures rich signals about user experiences, preferences, and concerns. However, the scale, redundancy, and imbalance of such corpora make it challenging to analyze opinions effectively, particularly when the goal is to generate summaries that remain faithful to the diversity of viewpoints expressed. This paper presents a framework that preserves semantics in LLM-based opinion summarization while minimizing token usage. We combine multidimensional classification (e.g., sentiment, topics) with a family of stratified sampling strategies to select compact yet representative subsets of opinions before prompting the LLM. Tailored prompts then produce balanced summaries that surface the salient aspects expressed in the opinions (e.g., strengthsand weaknesses of products/hotels) . Experiments on Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts demonstrate that our method significantly reduces token usage and computational cost while consistently outperforming traditional AI-based and standard LLM summarization baselines in terms of content coverage, balance, and semantic preservation.  \nKeywords: Large Language Models; Review Summarization; Opinion Mining  \n1 Introduction  \nUser-generated opinions in the form of product reviews, hotel assessments, social posts, and peer-support discussions constitute a rich source of information for capturing user experiences, preferences, and concerns [28, 31] . Organizations rely on these signals to evaluate product performance, identify service issues, and monitor emerging trends, while users consult them to gather information, compare alternatives, and make informed decisions [19] . Yet the scale, heterogeneity, and redundancy of these corpora make effective opinion analysis challenging, especially when one aims to preserve the diversity and nuance of the viewpoints expressed [22, 27] .  \nLarge Language Models (LLMs) offer strong capabilities for opinion classification and analysis, but applying them directly to full corpora is often inefficient and prone to bias: context windows are finite, the cost of processing large inputs is high, and model outputs tend to overweight majority viewpoints when the input distribution is not controlled [25] . Existing pipelines frequently rely on single-dimension analysis (e.g., sentiment only [25, 27]), fail to account for the need to balance multiple semantic facets such as topics and emotions, or lose important information when opinion sets are reduced without principled criteria. These limitations highlight the need for methods that preserve the diversity and structure of opinionated content while supporting efficient large-scale analysis.  \nThis paper introduces a framework designed to preserve semantics in LLM-based opinion summarization while minimizing input size. The approach first applies multidimensional classification—such as sentiment, topics, emotion, and optional domain-specific facets—to impose structure on the corpus. The framework then employs stratified sampling strategies to  \nselect compact yet representative subsets of opinions before prompting the LLM. By feeding the model a subset that is already balanced across facets and rich in informative content, the LLM can generate summaries that are both more faithful and less biased than those obtained from raw or arbitrarily reduced inputs. Tailored prompts guide the model to recover the salient aspects expressed in the corpus, such as product strengths and weaknesses, hotel service issues, and political support rationales. Unlike RAG-based approaches [10, 16] or methods that rely on","cbCaibcTBuheLT7U","https://ap.wps.com/l/cbCaibcTBuheLT7U","pdf",1152316,4,1,15,"English","en",105,"# Abstract\n# Introduction\n# Framework Overview\n## Multidimensional Classification\n## Stratified Sampling\n## Facet-Aware Prompting\n# Experiments and Results\n## Amazon Reviews\n## Tripadvisor Hotel Reviews\n## Twitter/X Discussions\n# Contributions and Extensions","[{\"question\":\"What problem does the paper address in opinion summarization with LLMs?\",\"answer\":\"Opinion corpora are large, redundant, and imbalanced, making it difficult to analyze opinions effectively and generate summaries that preserve the diversity and nuance of viewpoints. Directly feeding full corpora to LLMs is inefficient due to limited context windows and high processing cost.\"},{\"question\":\"How does the proposed framework reduce token usage while preserving semantics?\",\"answer\":\"It performs multidimensional classification (e.g., sentiment, topics, emotions) to structure the corpus, then applies stratified sampling to select a compact but representative subset that is balanced across semantic facets. The LLM then receives this subset using tailored prompts to surface salient aspects expressed in the opinions.\"},{\"question\":\"What datasets are used to evaluate the method, and what outcomes are reported?\",\"answer\":\"Experiments use Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts. The method significantly reduces token usage and computational cost while outperforming traditional AI-based and standard LLM summarization baselines in content coverage, balance, and semantic preservation.\"}]",1784212313,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"large-language-models-for-token-efficient-and-semantic-preserving-opinion-summarization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/large-language-models-for-token-efficient-and-semantic-preserving-opinion-summarization/86514/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in opinion summarization with LLMs?","Question",{"text":75,"@type":76},"Opinion corpora are large, redundant, and imbalanced, making it difficult to analyze opinions effectively and generate summaries that preserve the diversity and nuance of viewpoints. Directly feeding full corpora to LLMs is inefficient due to limited context windows and high processing cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework reduce token usage while preserving semantics?",{"text":80,"@type":76},"It performs multidimensional classification (e.g., sentiment, topics, emotions) to structure the corpus, then applies stratified sampling to select a compact but representative subset that is balanced across semantic facets. The LLM then receives this subset using tailored prompts to surface salient aspects expressed in the opinions.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets are used to evaluate the method, and what outcomes are reported?",{"text":84,"@type":76},"Experiments use Amazon product reviews, Tripadvisor hotel reviews, and X/Twitter posts. The method significantly reduces token usage and computational cost while outperforming traditional AI-based and standard LLM summarization baselines in content coverage, balance, and semantic preservation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]