[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85855-en":3,"doc-seo-85855-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85855,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Graph Representation of RaagBase: A Unique Dataset for Hindustani Music","Raag classification is a core MIR task for Hindustani Music and supports recommendation, education, archiving, and intelligent search. Existing work often relies on annotated audio or labeled datasets, yet complete note sequences better preserve temporal structure and musical context. RaagBase introduces a notation-based text dataset of note sequences from Pt. Bhatkhande. It also presents a graph-based raag representation that models dominance and note absence, enabling graph clustering to produce coherent groups aligned with ground-truth labels.","arXiv :2607 . 10229v 1 [ cs . SD] 11 Jul 2026  \nGRAPH REPRESENTATION OF RAAGBASE: A UNIQUE DATASET FOR  \nHINDUSTANI MUSIC  \nCHANDAN MISRA AND SWARUP CHATTOPADHYAY  \nAbstract . Raag classification is a fundamental MIR task for Hindustani Music, with applications in recommendation, education, archiving, and intelligent search. However, raag clustering remains underexplored, as most existing approaches rely on annotated audio or labeled datasets. While annotated melodic phrases capture characteristic patterns, complete note sequences preserve temporal structure and contextual dependencies, making them more suitable for data-driven modeling. In this work, we introduce RaagBase, a notation-based text dataset consisting of note sequences from compositions by Pt.  \nBhatkhande. Furthermore we propose a novel graph-based representation of raag structures by modeling the dominance and absence of notes in compositions. Each composition is represented as a node, and the edges between two compositions corresponds the similarities between them based on the note frequency distribution. Further, we apply established graph clustering techniques to identify groups of similar raag compositions. Experimental results demonstrate highly coherent clusters with strong agreement to ground-truth raag labels, thereby validating both the dataset and the proposed representation.  \nThe dataset is publicly available at [https://anonymous.4open.science/r/RaagBase-5427](https://anonymous.4open.science/r/RaagBase-5427) .  \n1. Introduction  \nHindustani Music, also known as North Indian Classical Music, is one of the two major art music traditions of the Indian subcontinent. It is characterized by two primary frameworks: Raag, the melodic framework, and Taal, the rhythmic framework. Raag can be viewed as lying between a scale and a tune [1], providing a structured grammar that defines characteristic melodic sequences known as Aroh (ascending) and Avroh (descending) . These sequences serve as the building blocks for constructing melodies that evoke specific moods and also provide cues for identifying raags.  \nIdentification of a raag in any composition finds applications and research in music learning [2], music information retrieval (MIR) [3–6], automatic classification of moods of compositions [7,8], recommending music to listeners, creating music generation systems, etc. Existing approaches to this problem primarily rely on annotated audio or music datasets [9–11] with labeled data primarily to train supervised models [12–14] . However, such datasets are predominantly offer audio-based features, and require substantial manual effort and expert intervention for their creation [13] . While these datasets include annotated melodic phrases that capture characteristic patterns of a raag, the lack of complete note sequences limits data-driven statistical analysis, such as note frequency distribution, n-gram modeling [15], Hidden Markov Model [16] and transition probability estimation [17] .  \nIn this work, we propose an alternative perspective by formulating raag identification as a graph clustering problem over musical compositions. Rather than directly assigning labels, each composition is represented through its note frequency distribution derived from notated music sheets. Although improvisation in Indian art music introduces considerable variability in note sequences through ornamentation, timing, and phrase expansion—even within the same raag and composition—these variations largely preserve the underlying tonal distribution of notes, making frequency-based representations robust to performance-specific differences. The intuition behind this framework is that compositions belonging to similar raags share comparable sets of permissible notes (i.e., aroh and avroh), leading to similar tonal distributions. Consequently, they are expected to exhibit strong similarity measures and naturally form coherent clusters. Each such distribution is represented as a node in ","cbCaibeq34xskUG2","https://ap.wps.com/l/cbCaibeq34xskUG2","pdf",1877917,3,1,"English","en",105,"# Introduction\n## Raag identification as graph clustering\n## RaagBase dataset overview","[{\"question\":\"How is each composition represented in the proposed graph model?\",\"answer\":\"Each composition is represented as a node, and edges reflect similarity between compositions based on note frequency distributions using measures such as cosine similarity and Euclidean distance.\"}]",1784206717,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"graph-representation-of-raagbase-a-unique-dataset-for-hindustani-music","",{"@graph":35,"@context":76},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/graph-representation-of-raagbase-a-unique-dataset-for-hindustani-music/85855/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"How is each composition represented in the proposed graph model?","Question",{"text":74,"@type":75},"Each composition is represented as a node, and edges reflect similarity between compositions based on note frequency distributions using measures such as cosine similarity and Euclidean distance.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]