[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86387-en":3,"doc-seo-86387-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86387,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Tokenization vs Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition","Inertial measurement unit (IMU)-based online handwriting recognition can process input gathered on different writing surfaces, yet it struggles with uneven character distributions and strong inter-writer variability. This study systematically compares subword tokenization and concatenation-based data augmentation on the OnHWWords500 dataset. Results show tokenization improves writer-independent generalization and reduces WER from 15.40% to 12.99%, while degrading on writer-dependent splits. Concatenation-based augmentation regularizes training, cutting CER by 34.5% and WER by 25.4%.","arXiv :2603 . 16883v2 [ cs .CV] 12 Jul 2026  \nTokenization vs. Augmentation: A Systematic Study of Writer Variance in IMU-Based Online Handwriting Recognition  \nJindong Li 1 ,4[0000−0002−3550−1660], Dario Zanca 1[0000−0001−5886−0597], Vincent Christlein 1[0000−0003−0455−3799], Tim Hamann2[0000−0003−3562−6882], Jens Barth2[0000−0003−3967−9578], Peter Kämpf2 , and Björn  \nEskofier 1 ,3 ,4 ,5[0000−0002−0417−0336]  \n1 Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany  \n2 STABILO International GmbH, Heroldsberg, Germany  \n3 Ludwig-Maximilians-Universität München, Munich, Germany  \n4 Munich Center for Machine Learning, Munich, Germany  \n5 Helmholtz Zentrum München-German Research Center for Environmental Health, Neuherberg, Germany  \nAbstract. Inertial measurement unit-based online handwriting recognition enables the recognition of input signals collected across different writing surfaces but remains challenged by uneven character distributions and inter-writer variability. In this work, we systematically investigate two strategies to address these issues: subword tokenization and concatenation-based data augmentation. Our experiments on the OnHWWords500 dataset reveal a clear dichotomy between handling inter-writer and intra-writer variance. On the writer-independent split, structural abstraction via Bigram tokenization significantly improves generalization to unseen writing styles, reducing the word error rate (WER) from 15.40 % to 12 .99 % . In contrast, on the writer-dependent split, tokenization degrades performance due to vocabulary distribution shifts between the training and validation sets. Instead, our proposed concatenation-based data augmentation acts as a powerful regularizer, reducing the character error rate by 34 .5 % and the WER by 25 .4 % . Further analysis shows that short, low-level tokens benefit model performance and that the performance gains from concatenation-based data augmentation surpass those achieved by proportionally extended training. These findings reveal a clear variance-dependent effect: subword tokenization primarily mitigates inter-writer stylistic variability, whereas concatenation-based data augmentation effectively compensates for intra-writer distributional sparsity.  \nCode is available at [https://github.com/jindongli24/TVA](https://github.com/jindongli24/TVA).  \nKeywords: Online Handwriting Recognition · Time-Series Analysis · Tokenization · Data Augmentation  \n1 Introduction  \nOnline Handwriting Recognition (OnHWR) has long been a cornerstone of natural user interfaces. It enables the digitization of human language through the analysis  \n2 J. Li et al.  \nof temporal stroke trajectories [7,3] . Unlike offline recognition, which processes static images of completed text, OnHWR leverages the dynamic sequence of writing to decode intent. Inertial Measurement Unit (IMU)-based methods extend this capability beyond touchscreens by employing wearable sensors to capture motion in untethered environments [21, 15] .  \nHowever, robust implementation faces two main challenges: uneven character distribution and high variability in writing styles. In natural languages, character frequency varies significantly. For example, vowels typically appear more often than consonants [20] . This skewed distribution limits the samples available for rare characters and leads to poor generalization. Furthermore, individual writing styles differ among users. The same shape can be interpreted as different characters depending on the writer [8,2] . Consequently, training a model that generalizeswell across diverse writing styles remains a difficult task.  \nTo address these challenges, data augmentation and tokenization are employed to enhance model stability and generalization. Data augmentation synthetically expands the training set and artificially balances the data distribution. By generating variations of infrequent characters through geometric transformations [1] or generative modeling [14], this ","cbCaidlEdrUiatPD","https://ap.wps.com/l/cbCaidlEdrUiatPD","pdf",354541,3,1,17,"English","en",105,"# Abstract\n# Introduction\n## Problem: character imbalance and writer variability\n## Related strategies: tokenization and augmentation\n# Contributions\n## Subword tokenization (Bigram, BPE, Unigram)\n## Concatenation-based data augmentation","[{\"question\":\"What two strategies are compared to address writer variance in IMU-based online handwriting recognition?\",\"answer\":\"The work compares subword tokenization and concatenation-based data augmentation. The goal is to mitigate inter-writer stylistic variability and intra-writer distributional sparsity under different variance regimes.\"},{\"question\":\"How does Bigram tokenization affect performance on the writer-independent split?\",\"answer\":\"Bigram tokenization improves generalization to unseen writing styles on the writer-independent split, reducing the word error rate (WER) from 15.40% to 12.99%.\"},{\"question\":\"Why does tokenization behave differently on the writer-dependent split, and what helps there instead?\",\"answer\":\"On the writer-dependent split, tokenization degrades performance due to vocabulary distribution shifts between training and validation. Concatenation-based data augmentation instead acts as a strong regularizer, reducing character error rate (CER) by 34.5% and WER by 25.4%.\"}]",1784211435,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tokenization-vs-augmentation-a-systematic-study-of-writer-variance-in-imu-based-online-handwriting-recognition","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/tokenization-vs-augmentation-a-systematic-study-of-writer-variance-in-imu-based-online-handwriting-recognition/86387/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What two strategies are compared to address writer variance in IMU-based online handwriting recognition?","Question",{"text":75,"@type":76},"The work compares subword tokenization and concatenation-based data augmentation. The goal is to mitigate inter-writer stylistic variability and intra-writer distributional sparsity under different variance regimes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Bigram tokenization affect performance on the writer-independent split?",{"text":80,"@type":76},"Bigram tokenization improves generalization to unseen writing styles on the writer-independent split, reducing the word error rate (WER) from 15.40% to 12.99%.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does tokenization behave differently on the writer-dependent split, and what helps there instead?",{"text":84,"@type":76},"On the writer-dependent split, tokenization degrades performance due to vocabulary distribution shifts between training and validation. Concatenation-based data augmentation instead acts as a strong regularizer, reducing character error rate (CER) by 34.5% and WER by 25.4%.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]