[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82059-en":3,"doc-seo-82059-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82059,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","THE BIO COLLECTION: Unified Pre-Training Scale LLM Corpus for Biology","Large language models for biology require training corpora that capture genuine biological understanding, yet existing resources are fragmented across heterogeneous formats and are not organized for language model pre-training. THE BIO COLLECTION provides a unified 52.6B-token corpus that converts scattered molecular, protein, genomic, single-cell, and pathway resources into training-ready instruction data. Each record is enriched with tool-computed biological properties and paired with THE BIO COLLECTION-E VAL for recognition, generation, and prediction. Training Gravity-16B-A3B on THE BIO COLLECTION more than doubles evaluation performance across all domains while preserving general language ability.","arXiv :2607 .08803v1 [ q-bio .QM] 9 Jul 2026  \nTHE BIO COLLECTION:  \nUnified Pre-Training Scale LLM Corpus for Biology  \nHyunjin Seo1,∗, Hyeon Hwang1,∗, Gyubok Lee2, Jay Shin1, Jimin Park3, Taesoo Kim4, Sanghoon Lee5, Hongjoon Ahn1,‡, Sungjun Han1,‡, Sangwon Jung1,‡  \n1 Trillion Labs 2 KAIST 3 SK Biopharmaceuticals Co., Ltd. 4 Lunit Inc. 5 AIGEN Sciences Inc.  \n∗ First Author  \n‡ Project Lead  \nAbstract  \nThe push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology. However, existing biological resources, such as molecular databases, protein repositories, genomic annotations, single-cell atlases, and pathway databases, are scattered across heterogeneous formats and remain unorganized into a cohesive corpus for language model training. We present THE BIO COLLECTION, a 52.6B-token pre-training-scale corpus that converts these disparate resources into a unified, training-ready form spanning small molecules, proteins, genomic sequences, cells, and pathways. Beyond consolidating existing data, THE BIO COLLECTION enriches each record with tool-computed biological properties and introduces new instruction tasks for capabilities that current corpora barely cover. We pair the corpus with THE BIO COLLECTION-E VAL, a matched suite probing recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings. Holding the base Gravity-16B-A3B architecture fixed, training on THE BIO COLLECTION more than doubles its overall score on THE BIO COLLECTION-E VAL with gains in every domain, while leaving general linguistic ability nearly intact.  \n Training corpus: THE BIO COLLECTION  \n Evaluation dataset: THE BIO COLLECTION-E VAL  \n Model: Gravity-bio-16B-A3B  \n1. Introduction  \nLarge language models (LLMs) for biology (BioLM) (Jang et al., 2026; Li et al., 2026; Wang et al., 2025a,b; Xia et al., 2025) aim to extend LLMs beyond general text understanding toward the representation and reasoning of biological systems and generation of biological entities. Unlike task-specific predictors trained for a single modality or endpoint in biology (Cui et al., 2024; Lin et al., 2023; Zhou et al., 2024), BioLM seek to operate over diverse biological entities and processes, including small molecules, proteins, genomic sequences, cell states, and pathways. Understanding this broader spectrum is important as many biological questions involve not only recognizing isolated entities, but also understanding their properties, functions, interactions, and effects across different biological contexts (Mazein et al., 2024; Mohamed et al., 2021; Simon et al., 2024) .  \nFigure 1 | The construction pipeline of THE BIO COLLECTION. We first integrate scattered biological resources and refine them into structured texts. Subsequently, we enrich the biological information in free-text stream with computational tools. Finally, we construct new instruction datasets for domains that are underrepresented in the corpora.  \nRealizing this vision requires a large-scale training corpus for biology that exposes the model to the breadth and structure of biology. The field has already produced many useful resources, including molecular databases, protein repositories, genomic annotations, single-cell atlases, pathway databases, and computational biology tools. However, these resources are rarely organized as a large-scale, LLM training-friendly data corpus for BioLM; instead, they are scattered across heterogeneous formats such as tables, sequence files, graph edges, structured annotations, or isolated instruction schemas (Cai et al., 2026; Richard et al., 2024; Seo et al., 2026; Shen et al., 2024; The Tabula Sapiens Consortium, 2022; Yu et al., 2024) . Nevertheless, no prior effort has consolidated these resources into a large-scale pre-training corpus that teaches BioLMs broad spectrum of biological knowledge spanning molecules, proteins, genomes,","cbCaiizUuNxtu1xM","https://ap.wps.com/l/cbCaiizUuNxtu1xM","pdf",821038,1,32,"English","en",105,"# Introduction\n## BioLM motivation and challenges\n## Corpus construction pipeline\n## Corpus and evaluation design\n## Training setup and evaluation results","[{\"question\":\"How is THE BIO COLLECTION-e VAL used to evaluate models?\",\"answer\":\"THE BIO COLLECTION-E VAL is a matched evaluation suite that probes recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings.\"}]",1784177885,81,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"the-bio-collection-unified-pre-training-scale-llm-corpus-for-biology","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-bio-collection-unified-pre-training-scale-llm-corpus-for-biology/82059/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How is THE BIO COLLECTION-e VAL used to evaluate models?","Question",{"text":75,"@type":76},"THE BIO COLLECTION-E VAL is a matched evaluation suite that probes recognition, generation, and prediction across molecular, protein, genomic, cellular, and cross-domain settings.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]