[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81807-en":3,"doc-seo-81807-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81807,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","The Wiola Architecture for Efficient Small Language Models","Wiola presents a clean-slate small language model architecture built from first principles with no structural lineage to GPT, LLaMA, Mistral, or Falcon. The design contributes five original components: Spiral Rotary Positional Encoding on a 3D helical manifold, Gated Cross-Layer Attention for inter-layer coherence, Adaptive Token Merging to reduce attention complexity while preserving length, DualStream Feed-Forward with learned per-dimension fusion, and WiolaRMSNorm to prevent representation collapse. Full mathematical derivations, comparisons, and four HuggingFace-compatible model sizes validate effectiveness.","The Wiola Architecture for Efficient Small  \nLanguage Models  \nAryuemaan Kumar Chowdhury  \nResearch and Development, Oscowl Ai IIT Hyderabad Hyderabad, India [reacharyu@oscowl.in](reacharyu@oscowl.in)  \nAfreen Shaik  \nResearch and Development, Oscowl Ai Hyderabad, India [afreen@oscowl.in](afreen@oscowl.in)  \nYaparla Bhargavi  \nResearch and Development, Oscowl Ai Hyderabad, India [bhargavi@oscowl.in](bhargavi@oscowl.in)  \narXiv :2607 .0 1394v 1 [ cs .AI] 1 Jul 2026  \nBrahma Kumar Research and Development, Oscowl Ai  \nHyderabad, India  \n[brahma@oscowl.in](brahma@oscowl.in)  \nAbstract—We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with any existing model family including GPT, LLaMA, Mistral, or Falcon. Wiola introduces five independently novel components: (i) Spiral Rotary Positional Encoding (SRPE), which embeds token positions on a three-dimensional helical manifold combining absolute, relative, and hierarchical positional signals; (ii) Gated Cross-Layer Attention (GCLA), providing each decoder layer with soft cross-attention access to compressed summaries of two preceding layers for inter-layer coherence; (iii) Adaptive Token Merging (ATM), which dynamically merges semantically redundant adjacent tokens in middle network layers to reduce attention complexity without information loss; (iv) DualStream Feed-Forward (DSFF), replacing the conventional MLP with two parallel streams fused by a learned per-dimension gate; and (v) WiolaRMSNorm, a modified normalisation introducing a per-dimension learned offset vector that prevents representation collapse. We provide complete mathematical derivations, architectural block diagrams, complexity analyses, and systematic comparisons against GPT-2, LLaMA-2, and Mistral. Wiola is released in four sizes (120M, 360M, 700M, and 1.5B parameters) and is fully compatible with the HuggingFace Transformers ecosystem, with all 22 architectural unit tests passing.  \nIndex Terms—small language model, novel architecture, spiral rotary positional encoding, gated cross-layer attention, adaptive token merging, transformer variant  \nI. INTRODUCTION  \nThe Transformer [1] has driven remarkable progress in natural language processing. Yet the dominant model families—GPT [2], LLaMA [4], Mistral [5], and their derivatives—share the same structural lineage with incremental differences in positional encoding or attention grouping. This conservatism leaves open fundamental architectural questions: Can a different positional geometry better capture multi-scale linguistic structure? Can inter-layer information routing improve longrange coherence in generated text? Can token-level redundancy be exploited to reduce quadratic attention cost?  \nWiola is a clean-slate SLM that addresses all three questions through five novel architectural components. Every sub  \nThis work was conducted as an independent research contribution. No external funding was received.  \ncomponent is derived from independent mathematical principles and verified to be structurally distinct from all prior published formulations.  \nThe primary contributions of this work are:  \n1) SRPE: A 3D helical positional encoding combining absolute, relative, and hierarchical position on a unified manifold with no extra parameters.  \n2) GCLA: Gated cross-layer attention providing inter-layer coherence via compressed layer summaries at negligible compute overhead.  \n3) ATM: Dynamic greedy token merging in middle layers reducing attention FLOPs by 5–9% during training with exact length restoration.  \n4) DSFF: A dual-stream parallel FFN with per-dimension learned fusion, separating local and global feature extraction.  \n5) WiolaRMSNorm: Modified RMS normalisation with per-dimension offset that counteracts representation collapse in deep stacks.  \n6) A production implementation with 22 passing unit tests and full HuggingFace Hub integration.  \nII. RELATED WORK  \nA. Positional Encodi","cbCaiv6jY7eLrlG0","https://ap.wps.com/l/cbCaiv6jY7eLrlG0","pdf",3126859,3,1,7,"English","en",105,"# Introduction\n## Primary Contributions\n# Related Work\n## Positional Encoding\n## Attention Variants\n## Feed-Forward Networks\n## Token Compression\n# Notation","[{\"question\":\"What is Wiola, and how is its architecture positioned relative to existing model families?\",\"answer\":\"Wiola is a fully original small language model architecture developed from first principles. It intentionally shares no structural lineage with major families such as GPT, LLaMA, Mistral, or Falcon.\"},{\"question\":\"Which five novel components does Wiola introduce?\",\"answer\":\"Wiola introduces Spiral Rotary Positional Encoding (SRPE), Gated Cross-Layer Attention (GCLA), Adaptive Token Merging (ATM), DualStream Feed-Forward (DSFF), and WiolaRMSNorm.\"},{\"question\":\"How does Wiola reduce attention cost while aiming to preserve information?\",\"answer\":\"Adaptive Token Merging (ATM) dynamically merges semantically redundant adjacent tokens in middle network layers, reducing attention complexity while restoring the exact sequence length.\"}]",1784176282,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-wiola-architecture-for-efficient-small-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/the-wiola-architecture-for-efficient-small-language-models/81807/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Wiola, and how is its architecture positioned relative to existing model families?","Question",{"text":75,"@type":76},"Wiola is a fully original small language model architecture developed from first principles. It intentionally shares no structural lineage with major families such as GPT, LLaMA, Mistral, or Falcon.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which five novel components does Wiola introduce?",{"text":80,"@type":76},"Wiola introduces Spiral Rotary Positional Encoding (SRPE), Gated Cross-Layer Attention (GCLA), Adaptive Token Merging (ATM), DualStream Feed-Forward (DSFF), and WiolaRMSNorm.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Wiola reduce attention cost while aiming to preserve information?",{"text":84,"@type":76},"Adaptive Token Merging (ATM) dynamically merges semantically redundant adjacent tokens in middle network layers, reducing attention complexity while restoring the exact sequence length.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]