[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84953-en":3,"doc-seo-84953-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84953,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System","Large language models (LLMs) demonstrate strong performance on linguistic tasks, yet response-quality evaluation still lacks breadth. Existing approaches often measure only one aspect, which fails to reflect the full range of model capabilities. This work proposes a multi-factor scoring paradigm that jointly assesses accuracy, conciseness, factual consistency, readability, and coherence, supported by a GUI for visual analysis. Experiments on TruthfulQA reveal notable strengths in reasoning while highlighting persistent weaknesses in complex facts and ambiguities.","Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System  \nYiming Gai12[0009-0002-0947-7182], Junde Lu12[0009-0003-3072-6195] , Xuefei  \nHuang2(􀀍)[0000−0002−8670−1283] and Ying Li12(􀀍) [0000-0002-1545-1811]  \n1 School of Computer Science and Engineering, Beihang University, Beijing, China  \n{gaiym,ljd2406107,[liying}@buaa.edu.cn](liying}@buaa.edu.cn)  \n2 Data Science and Intelligent Computing Laboratory, Hangzhou International Innovation Institute, Beihang University, Hangzhou, Zhejiang 311115,  \nP.R.China  \n{[xuefei.huang}@buaa.edu.cn](xuefei.huang}@buaa.edu.cn)  \nAbstract. The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs ’ strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.  \nKeywords: LLM Evaluation, Multi-factor Scoring, Large Language Models,  \nBenchmarking, Model Comparison  \n1 Introduction  \nLLMs have permeated various aspects of human life, from intelligent assistants to academic research, and their influence is ubiquitous. However, a scientific and comprehensive evaluation of LLM's strengths and weaknesses, particularly its performance across multiple testing criteria, remains a pressing challenge. Traditional evaluation methods have primarily focused on a single dimension, such as accuracy or fluency. However, with the diversification of application scenarios, relying on a single metric is no longer sufficient to fully capture the model's capabilities. As a result, the development of multi-dimensional evaluation frameworks has become a key research focus.  \n2 Yiming Gai , Junde Lu , Xuefei Huang , Ying Li  \nExisting research has proposed various methods for evaluating the quality of LLM responses. Metrics based on n-gram matching, such as BLEU and ROUGE, are widely used in machine translation and text generation tasks[1] . However, their limitations in semantic understanding and contextual coherence make them less suitable for open-domain question answering. Recently, semantic embedding-based evaluation methods have gained traction. BERTScore[2], for example, leverages pretrained models to compute semantic similarity between texts, addressing the shortcomings of traditional metrics. For factual consistency, datasets like TruthfulQA[3] have been developed to test a model ’s truthfulness and reasoning ability with specially designed questions. However, these methods often overlook user experience-related dimensions such as readability and conciseness, making it difficult to provide a comprehensive evaluation perspective.  \nIn specialized domains, such as healthcare or law, the evaluation requirements for LLMs become more complex. These domains not only demand high factual accuracy but also necessitate the proper use of technical terminology and logical coherence[4] . However, the applicability of existing evaluation methods has not been fully validated in these contexts, as traditional metrics and single-dimensional assessments are inadequate for addressing the diverse practical needs. To address this, this paper proposes a multi-factor scoring system that combines five key","cbCainI6JseHIr7n","https://ap.wps.com/l/cbCainI6JseHIr7n","pdf",1543906,1,12,"English","en",105,"# Introduction\n# Related work","[{\"question\":\"Why do existing LLM evaluation methods fall short?\",\"answer\":\"They often focus on a single dimension (e.g., accuracy or fluency), which cannot capture the full spectrum of capabilities required by diverse real-world scenarios.\"},{\"question\":\"What dimensions are included in the proposed multi-factor scoring system?\",\"answer\":\"The system scores five dimensions: accuracy, conciseness, factual consistency, readability, and coherence, and it also integrates ROUGE for evaluation support.\"},{\"question\":\"What do results on the TruthfulQA dataset show?\",\"answer\":\"The framework identifies mainstream LLM strengths in reasoning tasks (with a composite score reported up to 0.6104) while exposing common limitations when handling complex facts and ambiguous information.\"}]",1784199679,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"comprehensive-evaluation-of-large-language-model-responses-a-multi-factor-scoring-system","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/comprehensive-evaluation-of-large-language-model-responses-a-multi-factor-scoring-system/84953/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing LLM evaluation methods fall short?","Question",{"text":75,"@type":76},"They often focus on a single dimension (e.g., accuracy or fluency), which cannot capture the full spectrum of capabilities required by diverse real-world scenarios.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What dimensions are included in the proposed multi-factor scoring system?",{"text":80,"@type":76},"The system scores five dimensions: accuracy, conciseness, factual consistency, readability, and coherence, and it also integrates ROUGE for evaluation support.",{"name":82,"@type":73,"acceptedAnswer":83},"What do results on the TruthfulQA dataset show?",{"text":84,"@type":76},"The framework identifies mainstream LLM strengths in reasoning tasks (with a composite score reported up to 0.6104) while exposing common limitations when handling complex facts and ambiguous information.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]