[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127072-en":3,"doc-seo-127072-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127072,962084931830,"Theodore","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Identifying and Analyzing Reduplication Multiword Expressions in Hindi Text Using Machine Learning - SlideShare","The study addresses identifying and analyzing Reduplication Multiword Expressions (RMWEs) in Natural Language Processing for Hindi, aiming to extract repeated-word patterns and classify them into onomatopoeic, non-onomatopoeic, partial, and semantic types. A machine-learning-based identification method is proposed using linguistic patterns and statistical information, including a threshold boundary detection step for filtering. Semantic relations are analyzed with Jaccard distance and Sorensen Dice similarity, and the approach is evaluated on the public IITB Hindi corpus across consecutive thresholds.","Atul Mishra 1 , Alok Mishra 2  \n1 BML Munjal University, India  \n2 Faculty of Engineering, NTNU-Norwegian University of Science and Technology, Norway  \nAbstract - The task of identifying and analyzing Reduplication Multiword Expressions (RMWEs) in Natural Language Processing (NLP) involves extracting repeated words from various text forms and classifying them into Onomatopoeic, nonOnomatopoeic, partial, or semantic types. With the increasing use of low-resource languages in news, opinions, comments, hashtags, reviews, posts, and journals, this study proposes a machine learningbased RMWE identification method for Hindi text. The method employs linguistic patterns and statistical data, along with a proposed threshold boundary detection in statistical filtering. The Jaccard distance of dissimilarity and Sorensen Dice Coefficient of Similarity are used for semantic relation analysis. The proposed approach was evaluated using the publicly available Hindi corpus from IITB, measuring performance between two consecutive thresholds with the lowest error and highest recall. This study proposes an effective method for Indian computational linguistics, with experimental results highlighting its viability and utility, and providing a blueprint for current procedures.  \nKeywords - Linguistic patterns, natural language processing, computational linguistics, statistical data, threshold boundary detection.  \nDOI: 10.18421/TEM123-56  \n[https://doi.org/10.18421/TEM123-56](https://doi.org/10.18421/TEM123-56)  \nCorresponding author: Alok Mishra,  \nFaculty of Engineering, NTNU-Norwegian University of Science and Technology, Norway  \nEmail: [alok. mishra@ntnu. no](alok. mishra@ntnu. no)  \n[Received: 08 April 2023](Received: 08 April 2023).  \nRevised: 18 July 2023.  \nAccepted: 29 July 2023.  \nPublished: 28 August 2023.  \n1. Introduction  \nA multiword expression (MWE) is a lexeme (a basic lexical unit) composed of two or more separate lexemes e.g., रेल गाडी (“Rail gaadi”, Train), प्रधान मंत्री (“Pradhan Mantri”, Prime Minister) . Due to institutionalized usage, we tend to think of ‘रेल गाडी’ and ‘प्रधान मंत्री’ as a single concept. Here the concept crosses word boundaries. MWE are heterogeneous, treated as single words, unpredictable, non-literal translations that crosses word boundaries, and are restricted to sentence boundaries, i.e., MWEs can exist only within the sentence. In the proposed study, expressions within the sentence boundaries are considered. MWEs are made up of a few words (in the conventional sense), but they act as single words to some extent [1] . MWEs are important in applications like Machine Translation [2], Sentiment Analysis [3] and Information Retrieval systems [4] . Grammars define them inconsistently and are not sufficiently formalized in dictionaries or successfully extended to MT. MWE processing is therefore unpredictable, and non-literal.  \nReduplication [5] is a subcategory of MWE in which a string occurs in a repeated sequence, doubled, or several times within a larger syntactic unit in non-distinct positions. Reduplication means the repetition of units such as phonemes, morphology, word, phrase, clause, or utterances [6] . In this study, Hindi is chosen for research, which is part of the Indo-Aryan group within the Indo-Iranian branch of the Indo-European Language Family. Reduplication MWE in Hindi acts as expressions and phrases at the same time; sometimes in an unbendable design, and hence is  \nReduplication (Word Replication) is categorized as:  \n1. Onomatopoeic Expression: E.g., िटक िटक (Ṭik ṭik), खड़ खड़ (Khaḍa khaḍa), िझल िमल (jhil mil) -[meaning – a Sound]  \n2. Non–Onomatopoeic Expression: E.g., अभीअभी (just now, Abhi Abhi)  \n3. Partial Reduplication: E.g., खाना – वाना (to make the rhythm with food, Khana vaana), लाल वाल(to make the rhythm with color, Laal vaal)  \n4. Semantic Reduplication: E.g., िदन रात (Day night, Din raat), धन-दौलत (Wealth, Dhan Daulat) Table 1 shows how reduplication expressions in Hindi affe","cbCaipPC1TRWJjOp","https://ap.wps.com/l/cbCaipPC1TRWJjOp","pdf",486300,1,10,"English","en",105,"# Introduction\n## Multiword Expressions (MWEs)\n## Reduplication as an MWE Subcategory\n## Types of Reduplication in Hindi\n## Impact on Translation and NLP Tasks","[{\"question\":\"什么是 Reduplication Multiword Expressions（RMWEs）？\",\"answer\":\"RMWEs 是一种多词表达类别，特点是同一字符串在更大的句法单元中以重复序列形式出现，并在非同一位置重复。研究将其视为在句内边界内出现的可识别表达。\"},{\"question\":\"该研究如何识别与分类 RMWEs？\",\"answer\":\"方法结合语言学模式与统计信息，并引入阈值边界检测用于统计过滤。随后利用相似性/距离度量分析语义关系，并按类型归类为不同类别。\"},{\"question\":\"实验如何评估模型效果？\",\"answer\":\"实验在 IITB 提供的公开印地语语料上进行测试，通过两个连续阈值区间比较性能，重点考察最低错误与最高召回表现。\"}]","Identifying and Analyzing Reduplication Multiword Expressions in Hindi Text Using Machine Learning - SlideShare | PDF",1785936674,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"identifying-and-analyzing-reduplication-multiword-expressions-in-hindi-text-using-machine-learning-slideshare","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/identifying-and-analyzing-reduplication-multiword-expressions-in-hindi-text-using-machine-learning-slideshare/127072/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"什么是 Reduplication Multiword Expressions（RMWEs）？","Question",{"text":76,"@type":77},"RMWEs 是一种多词表达类别，特点是同一字符串在更大的句法单元中以重复序列形式出现，并在非同一位置重复。研究将其视为在句内边界内出现的可识别表达。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"该研究如何识别与分类 RMWEs？",{"text":81,"@type":77},"方法结合语言学模式与统计信息，并引入阈值边界检测用于统计过滤。随后利用相似性/距离度量分析语义关系，并按类型归类为不同类别。",{"name":83,"@type":74,"acceptedAnswer":84},"实验如何评估模型效果？",{"text":85,"@type":77},"实验在 IITB 提供的公开印地语语料上进行测试，通过两个连续阈值区间比较性能，重点考察最低错误与最高召回表现。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":21,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]