[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81601-en":3,"doc-seo-81601-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},81601,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",6,"Technology","QQ Language Metadata Toolkit for Multilingual NLP","Multilingual NLP research now spans hundreds or thousands of languages, making language metadata hard to manage and consistently report. QQ presents a versioned metadata toolkit and browser explorer that compiles language metadata into a graph covering varieties, scripts, regions, identifiers, names, and relations. Through a Python API, CLI, and web explorer, users normalize identifiers, retrieve metadata, traverse relations, discover which external resources contain a language, and generate reproducible reporting tables with FAIR-oriented practices.","QQ: A Language Metadata Toolkit for Multilingual NLP  \nWessel Poelman 1 and Yiyi Chen2 and Miryam de Lhoneux 1  \n1LAGOM·NLP, Department of Computer Science, KU Leuven  \n2AAU-NLP, Department of Computer Science, Aalborg University [wessel.poelman@kuleuven.be](wessel.poelman@kuleuven.be)  \narXiv :2603 .00620v2 [ cs .CL] 10 Jul 2026  \nAbstract  \nMultilingual NLP research increasingly involves hundreds or thousands of languages across different datasets. Managing, discovering, and reporting language metadata becomes a common hurdle at these scales. We present QQ, a metadata toolkit and browser explorer. QQ compiles language metadata sources into a graph of language varieties, scripts, regions, identifiers, names, and relations, and exposesit through a Python API, a command-line interface, and a browser-based explorer. Users can normalize identifiers, retrieve metadata, traverse relations, and discover which external resources contain a language. We demonstrate QQ on three workflows: an audit of the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables. QQ supports FAIR-oriented metadata practices through versioning, open formats, and reusable interfaces.  \n1  \n􀂌 Explorer & PyPI  \nIntroduction  \n􀂇 Source u Demo  \nThe number of languages considered in multilingual NLP research has increased drastically in recent years (e.g., Adebara et al., 2023 ; ImaniGooghari et al., 2023) . Datasets, benchmarks, and models now regularly cover hundreds or thousands of languages. While good news for the field, this also creates a practical problem: language metadata becomes harder to manage consistently.  \nAdditionally, multilingual models are increasingly evaluated on more tasks and in more languages. As more datasets are combined across training and evaluation, a common problem becomes more apparent: datasets do not stick to one convention when using language identifiers. One dataset might list German as de, another as deu, and a third as stan1295 . This is easy to handle when dealing with ten or twenty languages,  \nbut with hundreds or thousands of languages, with several identifier types in active use, it becomes cumbersome at best and unmanageable at worst. This issue is amplified since languages need to be tracked across exploration, data cleanup, analysis, and reporting.  \nIdentifiers are only part of the problem. Language metadata matters a posteriori when reporting which languages were used in a study (Bender, 2019), and it matters a priori when choosing languages based on scripts, geography, families, or other properties. Currently, managing this often means using local mapping files, scattered scripts, and manual browsing across several metadata sources, such as Wikipedia or Glottolog (Hammarström et al., 2024) . As the number of languages grows, one-off scripts become increasingly brittle: the same language must be resolved repeatedly across different resources and experimental steps.  \nComing back to our German example: OPUS (Tiedemann, 2012) and Universal Dependencies (Nivre et al., 2020) use the ISO 639-1 code de; WMT uses the ISO 639-3 code deu ; older datasets might use ISO 639-2 ger ; some datasets, such as NLLB (Costa-jussà et al., 2024) and FineWeb-2 (Penedo et al., 2025), use tags such as deu_ Latn ; linguistic datasets often use the Glottocode stan1295 (Forkel and Hammarström, 2022); and Wikidata uses Q188 . These identifiers refer to the same language, but when combining datasets, they first have to be aligned.  \nThe same problem appears when trying to find datasets for a particular language. On the HuggingFace Hub (Lhoest et al., 2021), the language filter shows separate entries for what is effectively the same language, such as el and ell for Modern Greek. Some commonly-used linguistic resources, like WALS (Dryer and Haspelmath, 2013), invent their own identifiers that do not belong to any other standard (grk for Modern Greek) . This makes the discove","cbCaicJ2tjoC1Lgk","https://ap.wps.com/l/cbCaicJ2tjoC1Lgk","pdf",469955,1,11,"English","en",105,"# Introduction\n## Motivation: Identifier inconsistency at scale\n## Language metadata for discovery and reporting\n## How QQ addresses these issues\n# Contributions and Toolkit Overview\n## Interfaces: Python, CLI, Browser explorer\n## FAIR-oriented metadata practices","[{\"question\":\"What problem does QQ address in multilingual NLP research?\",\"answer\":\"QQ addresses the difficulty of managing and unifying language metadata when research spans hundreds or thousands of languages and datasets use inconsistent identifier conventions.\"},{\"question\":\"How does QQ represent language information?\",\"answer\":\"QQ compiles multiple language metadata sources into a versioned unified database and organizes them as a metadata graph of language varieties, scripts, regions, identifiers, names, and relations.\"},{\"question\":\"Which workflows does QQ support and how?\",\"answer\":\"QQ supports practical workflows such as auditing language tags on resources like the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables via its Python API, command-line interface, and browser explorer.\"}]","QQ Language Metadata Toolkit for Multilingual NLP | PDF",1784174647,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"qq-language-metadata-toolkit-for-multilingual-nlp","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/qq-language-metadata-toolkit-for-multilingual-nlp/81601/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does QQ address in multilingual NLP research?","Question",{"text":76,"@type":77},"QQ addresses the difficulty of managing and unifying language metadata when research spans hundreds or thousands of languages and datasets use inconsistent identifier conventions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does QQ represent language information?",{"text":81,"@type":77},"QQ compiles multiple language metadata sources into a versioned unified database and organizes them as a metadata graph of language varieties, scripts, regions, identifiers, names, and relations.",{"name":83,"@type":74,"acceptedAnswer":84},"Which workflows does QQ support and how?",{"text":85,"@type":77},"QQ supports practical workflows such as auditing language tags on resources like the HuggingFace Hub, linking resources that use different identifier systems, and generating reproducible language-reporting tables via its Python API, command-line interface, and browser explorer.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,114,119,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},8,"Research & Report",30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]