[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82574-en":3,"doc-seo-82574-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82574,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Svarna: An Open Corpus Workbench for Modern Greek","Svarna is a free, open-source, web-based corpus workbench designed for Modern Greek and built to solve the long-standing fragmentation of Greek language resources. It unifies five independently searchable databases—institutional, literary/developmental, dialectal, social media, and historical/folkloric—covering over 507M words and about 29M sentences. The platform offers KWIC concordancing, frequency and normalization, collocation extraction, discourse-marker distributions, n-grams and networks, register comparison, regex search, and optional LLM-driven pragmatic annotation. System components use SQLite FTS5 with a FastAPI backend, containerized deployment on Azure, MIT-licensed code, and straightforward user corpus ingestion and instance deployment.","arXiv :2607 .00970v 5 [ cs .CL] 9 Jul 2026  \nSvarna: An Open Corpus Workbench for Modern Greek  \nStergios Chatzikyriakidis  \nDepartment of Philology, University of Crete, Rethymno, Greece [stergios. chatzikyriakidis@gu. se](stergios. chatzikyriakidis@gu. se)  \nAbstract  \nThis paper introduces Svarna, a free, open-source, web-based corpus workbench for Modern Greek. Svarna integrates five databases covering various registers, institutional, literary, dialectal, social media, and historical, to provide a total of more than 507 million words and around 29 million sentences. This platform attempts to address a chronic problem in Greek language technology. Although various corpus resources exist, they are scattered across different platforms, and in many cases, institutional access is restricted or they are no longer available online. Svarna integrates a number of these resources into a single interface that can be used without logging in, installation, or specialized training. SVARNA provides a concordancer with KWIC marking capabilities, frequency analysis including register-by-register normalization, collocation extraction using mutual information, a dictionary of 93 Greek discourse markers providing distribution profiles, text-level analysis tools including n-grams, varieties, and collocation networks, register comparison using log-ratio, regular expression search, and an optional LLM layer for pragmatic annotation and free research mode. This platform is built upon SQLite FTS5 full-text indices provided via a FastAPI backend, deployed as Docker containers on Azure, and released under the MIT license. Source code, build scripts, and deployment configurations are publicly available on GitHub. What is important, is that users can easily ingest their own corpora and deploy their own instances. This paper describes the system design, corpus structure, and use cases demonstrating the various queries supported by the platform. Svarna serves as the first step in exploring available data and aspires to lay the foundation for more comprehensive research in the future.  \nKeywords: Modern Greek; corpus linguistics; concordancer; open-source; language resources; discourse markers  \n1 Introduction  \nModern Greek is spoken by approximately 13 million people and possesses a written tradition that has been continuously used for centuries. However, the Greek corpus infrastructure is surprisingly fragmented. Researchers, students, and language experts who wish to search, query, and analyze Greek texts on a large scale face difficulties due to inaccessible and scattered materials.  \nOf course, materials are not entirely nonexistent. Over the past 20 years, valuable resources on the Greek language have been produced through various corpus construction projects. The Hellenic National Corpus (HNC) by ILSP/Athena RC was one of the early large-scale projects (Hatzigeorgiu et al., 2000; Mikros et al., 2005) . The Corpus of Greek Texts (CGT) built by Goutsos (2010) provides a balanced reference corpus. The CLARIN:EL infrastructure (Gavriilidou et al., 2023) provides various language resources through national portals. Greek data is included in large-scale  \nmultilingual datasets such as CC-100 (Conneau et al., 2020), OSCAR (Ortiz Suárez et al. , 2019), and mC4 (Xue et al., 2021) . Similar materials are also available in Europarl (Koehn, 2005) and OpenSubtitles (Lison and Tiedemann, 2016) . The Leipzig Corpora Collection contains Greek web data (Goldhahn et al., 2012), and Greek treebanks exist in Universal Dependencies (Prokopidis and Papageorgiou, 2017; Nivre et al., 2016) . Parliamentary minutes have been constructed as a dataset (Dritsa et al., 2022), and recently, dialect data has been collected from GRDD and GRDD+ datasets (Chatzikyriakidis et al., 2023, 2026) .  \nThe problem is not a lack of resources, but rather the difficulty in utilizing existing resources together. Access to some corpora is restricted to institutional credentials that are not ","cbCaibEkam76IyO1","https://ap.wps.com/l/cbCaibEkam76IyO1","pdf",1326497,4,1,17,"English","en",105,"# Introduction\n## Problem: fragmented Greek corpus infrastructure\n## Svarna overview and integrated corpora\n## Free, open deployment and public code","[{\"question\":\"What core analysis and search features does Svarna provide?\",\"answer\":\"Svarna includes a concordancer with KWIC marking, frequency analysis with register normalization, collocation extraction via mutual information, a discourse-marker dictionary with distribution profiles, and text-level tools such as n-grams and collocation networks. It also supports register comparison, regular-expression search, and optional LLM-based pragmatic annotation.\"}]",1784181598,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"svarna-an-open-corpus-workbench-for-modern-greek","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/svarna-an-open-corpus-workbench-for-modern-greek/82574/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What core analysis and search features does Svarna provide?","Question",{"text":75,"@type":76},"Svarna includes a concordancer with KWIC marking, frequency analysis with register normalization, collocation extraction via mutual information, a discourse-marker dictionary with distribution profiles, and text-level tools such as n-grams and collocation networks. It also supports register comparison, regular-expression search, and optional LLM-based pragmatic annotation.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]