[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119166-en":3,"doc-seo-119166-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119166,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","An Empirical Study on Document Similarity Comparison Evaluation Between Machine Learning Techniques and Human Experts - Evaluation metric and vision propagation measurement (VPMS) - Machine Learning vs humans","Focuses on document similarity comparison beyond pure accuracy in machine learning. Examines how machine-learning decision behavior aligns with human judgment across different domains, and evaluates whether organizational vision is effectively propagated to lower organizational levels. Uses word representation methods including count-based models (Bag of Words, TF-IDF), ANN models (Word2Vec, GloVe), and a VPMS approach that combines two methods. Compares model outputs and expert-group results from multiple perspectives, proposing a VPMS model that outperforms existing methods. Results support evaluation of vision dissemination and can inform text-based algorithms like customer segmentation for target marketing.","ISSN 1330-3651 (Print), ISSN 1848-6339 (Online) [https://doi.org/10.17559/TV-20231011001013](https://doi.org/10.17559/TV-20231011001013)  \nOriginal scientific paper  \nAn Empirical Study on Document Similarity Comparison Evaluation Between Machine Learning  \nTechniques and Human Experts  \nWon-Jung JANG  \nAbstract: Current machine-learning training focuses solely on accuracy. In this study, the weights of other dimensions were examined rather than measuring only the accuracy of machine learning. By comparatively analyzing the decision-making of machine learning and humans in various fields, this study examines how well organizational vision is propagated to lower levels of the organization. Also, the results evaluated by humans and machine learning models were comparatively analyzed from multiple perspectives. As numerical representation methods of words, count-based models (Bag of Words, TF-IDF), artificial neural network (ANN) models (Word2Vec, GloVe), anda vision propagation measurement (VPMS) model combining two methods were used to calculate the similarity between documents, which are comparatively analyzed with the actual results measured by an expert group. The findings of this study can be used as an evaluation metric for how effectively the vision of the upper organization is being disseminated to the lower-level organizations. Additionally, it could be utilized in developing algorithms such as customer segmentation for target marketing using textdata. The study makes two key contributions- (i) providing an extensive empirical comparison of document similarity analysis by different ML techniques versus human experts, and (ii) proposing a new VPMS model that outperforms existing methods.  \nKeywords: ANN model; count-based model; document similarity; ensemble learning model; machine learning  \n1 INTRODUCTION  \nIn the era of the 4IR (Fourth Industrial Revolution), personalized customer experience through big data-based value enhancement can have a significant impact on corporate profits. Data generation is expected to increase significantly owing to advancements in communication technology, such as the Internet and smartphones. According to an IDC report, 16.1 ZB (zettabyte) of digital data generated in 2016 will increase by 10 times in less than 10 years, while the rate of increase will also go up [9] . The Ministry of Science and ICT of South Korea reported that the data market size was 14 trillion won in 2018 andwill expand to 30 trillion won by 2023 based on relevant policies and programs that foster global competitiveness in related industries, such as data and artificial intelligence (AI) . The enormous net worth of Facebook is attributable to the large and unique data assets of the company [29] . Most companies are focusing on strengthening their AI-based products and service competitiveness, and securing big data based on AI has become crucial to improving competitiveness. In particular, increasing the competitiveness of a company using unstructured data is important because of the large volumes of data generated. From the technological advances in personal computers (PCs) in the 1990s to those in the Internet (2000) and smartphones (2007), digitized information and objects have become interconnected. In the digital era, the success of a company is determined by the efficiency of these connections. Companies such as Yahoo and Google are typical examples of these success stories. Therefore, searching for accurate information easily is critical for the survival of companies in the digital era. Finding an entity that quickly and easily provides information with a high level of similarity is an important issue. Because individuals communicate and connect with each other using documents, which are typical text data and may also be used for important decision-making, research based on documents is emerging as a significant research topic. Studies on developing algorithms to quickly and easily find accurate documents have","cbCaiepX8A4SK09z","https://ap.wps.com/l/cbCaiepX8A4SK09z","pdf",1892553,1,12,"English","en",105,"# Introduction\n## Background and motivation in the 4IR era\n## Document similarity as an evaluation indicator\n## Human-AI cooperation and the limits of accuracy-only training\n# Methodology (models and representations)\n## Count-based representations: Bag of Words and TF-IDF\n## ANN representations: Word2Vec and GloVe\n## VPMS model combining methods\n# Evaluation\n## Expert-group comparison\n## Multi-perspective assessment of similarity decisions\n# Results and contributions\n## Empirical comparison across ML techniques vs experts\n## VPMS model performance and implications\n# Applications\n## Vision propagation evaluation\n## Text-based targeting and customer segmentation","[{\"question\":\"What problem does the study address regarding current machine learning training?\",\"answer\":\"It argues that many machine-learning efforts optimize only accuracy, without examining other dimensions that shape decision behavior and alignment with human judgment.\"},{\"question\":\"Which document similarity representations and models are used in the study?\",\"answer\":\"The study uses count-based models (Bag of Words, TF-IDF), ANN-based models (Word2Vec, GloVe), and a VPMS model that combines two methods to compute document similarity.\"},{\"question\":\"How are the machine-learning results validated in the paper?\",\"answer\":\"Model outputs are comparatively analyzed against results measured by an expert group, and the comparison is evaluated from multiple perspectives.\"}]","An Empirical Study on Document Similarity Comparison Evaluation Between Machine Learning Techniques and Human Experts - Evaluation metric and vision propagation measurement (VPMS) - Machine Learning vs humans | PDF",1785722873,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"an-empirical-study-on-document-similarity-comparison-evaluation-between-machine-learning-techniques-and-human-experts-evaluation-metric-and-vision-propagation-measurement-vpms-machine-learning-vs-humans","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/an-empirical-study-on-document-similarity-comparison-evaluation-between-machine-learning-techniques-and-human-experts-evaluation-metric-and-vision-propagation-measurement-vpms-machine-learning-vs-humans/119166/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the study address regarding current machine learning training?","Question",{"text":76,"@type":77},"It argues that many machine-learning efforts optimize only accuracy, without examining other dimensions that shape decision behavior and alignment with human judgment.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which document similarity representations and models are used in the study?",{"text":81,"@type":77},"The study uses count-based models (Bag of Words, TF-IDF), ANN-based models (Word2Vec, GloVe), and a VPMS model that combines two methods to compute document similarity.",{"name":83,"@type":74,"acceptedAnswer":84},"How are the machine-learning results validated in the paper?",{"text":85,"@type":77},"Model outputs are comparatively analyzed against results measured by an expert group, and the comparison is evaluated from multiple perspectives.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]