[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82983-en":3,"doc-seo-82983-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82983,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES","Every chemical language model reading SMILES depends on tokenization, yet chemistry work has largely inherited byte-pair encoding (BPE) from natural language without testing alternatives. This preprint compares BPE and Unigram-LM under matched conditions using a fixed 165-token chemistry base across three corpus typologies and two boundary policies, varying vocabulary size where embeddings are learnable. The algorithms never converge: they form near-disjoint subword vocabularies, Unigram-LM yields 29–41% more tokens, and BPE is a strict coarsening of Unigram-LM on most molecules. Released trained tokenizers and measurements.","arXiv :2607 .0569 1v 1 [ cs .CL] 6 Jul 2026  \nWHERE TO CUT, HOW DEEP: BPE AND UNIGRAM-LM ON  \nCHEMISTRY SMILES  \nA PREPRINT  \nHunter Heidenreich  \nIndependent Researcher  \n[hheiden0@gmail.com](hheiden0@gmail.com)  \n[orcid.org/0009-0001-0335-4803](orcid.org/0009-0001-0335-4803)  \nABSTRACT  \nEvery chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE’s principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open: the complete glyph base already covers every conformant molecule, so the learned pieces add compression rather than coverage, and a tiny alphabet under hard valence constraints could drive two frequency-based algorithms to converge. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small vocabulary sizes where token embeddings are learnable, across three corpus typologies (diverse, drug-like, natural-products) and both pre-tokenization boundary policies. The two do not converge.  \nIn all 22 matched conditions they build near-disjoint subword vocabularies: cross-algorithm Jaccard overlap on the learned pieces above the shared base never exceeds 0. 161, and at most 0.05 once weighted toward the high-frequency pieces a model updates most. Unigram-LM also segments held-out molecules into 29–41% more tokens; the arms largely agree on where to cut but not how deeply, so BPE’s segmentation is a strict coarsening of Unigram-LM’s on 80–99% of molecules.  \nThe separation holds across corpus, boundary, and vocabulary size, persisting even at eight times that scale, past where embeddings remain learnable; only token-frequency imbalance attenuates in magnitude, shrinking with vocabulary size and most on the natural-products corpus, without closing.  \nThe subword algorithm is therefore a modeling decision, not a free default. We release all trained tokenizers and per-condition measurements.  \n1 Introduction  \nEvery chemical language model that operates on a SMILES string [53] begins with a tokenizer that maps it to integer IDs, fixing the vocabulary the model embeds, the granularity at which it operates, and hence the effective sequence length. It is a foundational design choice [2, 16], almost always made by default: SMILES tokenizers inherit byte-pair encoding (BPE) [43] from natural-language practice, and its chemistry descendants (SPE, APE, Smirk-GPE) are all BPE variants. The principal alternative, Unigram-LM [19], sees occasional chemistry use but has never been compared to BPE at matched conditions on a fixed chemistry-grammatical base.  \nThat choice was long overshadowed by a more basic axis, vocabulary coverage: whether every string gets a token or part falls through to [UNK] . Wadell et al. [51] found that coverage, not the subword scheme, separates chemistry tokenizers downstream, but only across heterogeneous bases at large native vocabularies, never isolating the subword-algorithm axis. Smirk [51] closes coverage with a complete 165-token OpenSMILES base [15],1 emitting no [UNK] on conformant input. Our study begins there: pushing the vocabulary into the small regime, we isolate the subword algorithm, still inherited from natural language, as the design choice under study.  \n1Following Wadell et al. [51]: 158 OpenSMILES glyphs plus seven special tokens (including [UNK]) . We write “glyph” for the 158 chemistry-grammatical pieces.  \nSerotonin Nicotine  \nSmirk-GPE (BPE)  \n8 tokens  \nUnigram-LM  \n13 tokens (+5)  \n13 tokens 20 tokens (+7)  \npiece size  \nsmaller  larger  \nFigure 1: Nicotine and serotonin under the two algorithms’V =1024 vocabularies (molecules in rows, algorithms in columns) . Hue marks the algorithm (blue Smirk-GPE/BPE, orange Unigram-LM); darker shading marks a larger piece, on the same scale for both. BPE builds a few large pieces spanning whole","cbCaik7rQEfwHpbi","https://ap.wps.com/l/cbCaik7rQEfwHpbi","pdf",1540661,4,1,42,"English","en",105,"# Introduction\n## Tokenization as a design choice\n## Motivation and convergence question\n## Study scope and key findings","[{\"question\":\"What problem does the preprint address about SMILES tokenizers?\",\"answer\":\"SMILES-based chemical language models rely on tokenizers that map strings to integer IDs, which fixes vocabulary, granularity, and effective sequence length. The work questions whether the commonly used BPE tokenization is actually appropriate for chemistry, compared with Unigram-LM.\"},{\"question\":\"How is the comparison between BPE and Unigram-LM conducted?\",\"answer\":\"The study keeps the chemistry base fixed (a complete 165-token setting) while comparing BPE and Unigram-LM across three corpus typologies and two pre-tokenization boundary policies, sweeping small vocabulary sizes where token embeddings are learnable.\"},{\"question\":\"What are the main findings about convergence and token segmentation?\",\"answer\":\"Across all matched conditions, the tokenizers do not converge: they build near-disjoint subword vocabularies, with cross-algorithm overlap staying extremely low. Unigram-LM segments held-out molecules into 29–41% more tokens, while BPE coarsens Unigram-LM’s segmentation on most molecules.\"}]",1784184445,106,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"where-to-cut-how-deep-bpe-and-unigram-lm-on-chemistry-smiles","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/where-to-cut-how-deep-bpe-and-unigram-lm-on-chemistry-smiles/82983/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the preprint address about SMILES tokenizers?","Question",{"text":75,"@type":76},"SMILES-based chemical language models rely on tokenizers that map strings to integer IDs, which fixes vocabulary, granularity, and effective sequence length. The work questions whether the commonly used BPE tokenization is actually appropriate for chemistry, compared with Unigram-LM.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the comparison between BPE and Unigram-LM conducted?",{"text":80,"@type":76},"The study keeps the chemistry base fixed (a complete 165-token setting) while comparing BPE and Unigram-LM across three corpus typologies and two pre-tokenization boundary policies, sweeping small vocabulary sizes where token embeddings are learnable.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main findings about convergence and token segmentation?",{"text":84,"@type":76},"Across all matched conditions, the tokenizers do not converge: they build near-disjoint subword vocabularies, with cross-algorithm overlap staying extremely low. Unigram-LM segments held-out molecules into 29–41% more tokens, while BPE coarsens Unigram-LM’s segmentation on most molecules.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]