[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-203787-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-203787-en":131},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","domain-fine-tuning-finbert-on-finnish-histopathological-reports-train-time-signals-and-downstream-correlations","Domain Fine-Tuning FinBERT on Finnish Histopathological Reports - Train-Time Signals and Downstream Correlations","","Domain fine-tuning of transformer models on unlabeled data is a common strategy when labeled examples are scarce. This paper investigates fine-tuning the Finnish BERT model on Finnish medical text and reports observations from the training process. It also evaluates whether downstream benefits of domain-specific pre-training can be predicted by analyzing embedding-geometry changes induced by fine-tuning. The focus addresses healthcare-AI scenarios where dataset and label acquisition is delayed.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/domain-fine-tuning-finbert-on-finnish-histopathological-reports-train-time-signals-and-downstream-correlations/203787/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/domain-fine-tuning-finbert-on-finnish-histopathological-reports-train-time-signals-and-downstream-correlations/203787.png","ImageObject",300,407,{"name":42,"@type":43},"Evangeline","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-07","2026-09-04",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",11,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"Why is domain fine-tuning useful in healthcare NLP with limited labels?","Question",{"text":63,"@type":64},"When labeled data is delayed or difficult to obtain, domain fine-tuning on unlabeled text can be performed first. The paper motivates this approach by typical healthcare-AI dataset and labeling constraints.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"What two aims does the paper pursue?",{"text":68,"@type":64},"The study aims to (1) describe observations from fine-tuning Finnish BERT on Finnish medical text and (2) attempt to predict downstream benefits from domain-specific pre-training by examining embedding changes caused by fine-tuning.",{"name":70,"@type":61,"acceptedAnswer":71},"How are train-time observations connected to downstream classification performance?",{"text":72,"@type":64},"The paper uses training-time signals (including changes reflected in embeddings and loss behavior) to look for correlations with classification task performance. It reports tentative correlations and notes that replication is needed.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},203787,1788563882,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,111,115,120,123,127],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":112,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":113,"slug":114},8,30,"research-report",{"id":116,"doc_module":4,"doc_module_name":25,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":25,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":25,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":25,"category_name":129,"show_sort_weight":97,"slug":130},19,"General","general",{"code":4,"msg":82,"data":132},{"doc_id":79,"user_id":133,"nickname":42,"user_avatar":134,"doc_module":4,"category_id":112,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":118,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":104},13056703019662,"https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188","arXiv :2604 . 14815v1 [ cs .CL] 16 Apr 2026  \nDomain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations  \nRami Luisto 1,2,3,*, Liisa Pet¨ainen 1 , Tommi Gr¨onholm 1 , Jan B¨ohm4 , Maarit Ahtiainen4 , Tomi Lilja4 , Ilkka P¨ol¨onen 1 ,  \nSami ¨Ayr¨am¨o 1, 5  \n1 Faculty of Information Technology, University of Jyv¨askyl¨a, Jyv¨askyl¨a, Finland.  \n2 Digital Workforce Services, Helsinki, Finland.  \n3 Heart and Lung Center, Helsinki University Hospital, Helsinki, Finland.  \n4 Central Finland Biobank, Jyv¨askyl¨a, Finland.  \n5Wellbeing Services County of Central Finland, Jyv¨askyl¨a, Finland.  \n*[Corresponding author. E-mail:](Corresponding author. E-mail: rami.m.luisto@jyu.fi)[ rami.m.luisto@jyu.fi](Corresponding author. E-mail: rami.m.luisto@jyu.fi);  \nAbstract  \nIn NLP classification tasks where little labeled data exists, domain fine-tuning of transformer models on unlabeled data is an established approach. In this paper we have two aims. (1) We describe our observations from fine-tuning the Finnish BERT model on Finnish medical text data. (2) We report on our attempts to predict the benefit of domain-specific pre-training of Finnish BERT from observing the geometry of embedding changes due to domain fine-tuning. Our driving motivation is the common† situation in healthcare AI where we might experience long delays in acquiring datasets, especially with respect to labels.  \n1 Introduction  \nEver since ULMFiT [1] the idea of separating the training of a language model to different stages of specialization (typically pre-training, domain fine-tuning and task-specific  \n†First author’s professional anecdote.  \n1  \ntraining) has been a core tenet in any modern NLP. This approach is particularly suitable for models based on the transformer architecture, see e.g. [2–5] . Recent work has continued to study the specific benefits of domain fine-tuning, e.g. in the catchily titled Don’t stop pretraining: Adapt language models to domains and tasks [6] study how Domain-adaptive Pre-Training (DPT) compares to Task-Adaptive Pre-Training (TAPT) 1 .  \nEspecially in the realm of Large Language Models (LLMs) it can be crucial to predict from smaller exploratory training runs what would be the optimal configurations of model weights, train data amount and compute budget. The literature records various power laws that can be used to extrapolate model training behaviour. Such power laws exist both for comparing models within a data domain, see e.g. [7, 8] and the references within, or to compare behaviour across domains, see [9] . Though such work is typically focused on LLMs instead of Small Language Models (SLMs) like BERT, the underlying transformer architecture is the same, and many of the effects transfer.  \nMLM Training Dynamics with Epoch Markers  \n(Circles mark epoch boundaries)  \nLoss  \n3.0  \n2.5  \n2.0  \n1.5  \n1.0  \n0 50 100 150 200 250 300  \nTokens (millions)  \nFig. 1: Observations on the train-time loss of FinBERT on various datasets it is more or less “familiar” with. Histopathological dataset shows massive changes in loss.  \nIn the current work our aim is twofold. First we study how we might observe the effect of DFT on a dataset where we do not have training labels available. Second we try to replicate and observe the benefits that DFT have towards a classification task, with an eye on DFT training time observations that might correlate with classification task performance. We emphasize that our aim is to observe the phenomenon and make qualitative assessments. In particular, we do not aim for optimal classification algorithms as we wish to observe performance differentials rather than absolute levels.  \n1 In this paper we do not make a differentiation at this resolution, and just talk of Domain Fine-Tuning, DFT.  \n2  \nOur driving motivation here is NLP work in the healthcare sector, in particular with regards to Finnish medical text data. Finnish is a minority language with complex g","cbCaihkfLRP9Qc1S","https://ap.wps.com/l/cbCaihkfLRP9Qc1S","pdf",940456,"English","# Abstract\n# Introduction\n## Domain adaptation in low-resource medical languages","[{\"question\":\"Why is domain fine-tuning useful in healthcare NLP with limited labels?\",\"answer\":\"When labeled data is delayed or difficult to obtain, domain fine-tuning on unlabeled text can be performed first. The paper motivates this approach by typical healthcare-AI dataset and labeling constraints.\"},{\"question\":\"What two aims does the paper pursue?\",\"answer\":\"The study aims to (1) describe observations from fine-tuning Finnish BERT on Finnish medical text and (2) attempt to predict downstream benefits from domain-specific pre-training by examining embedding changes caused by fine-tuning.\"},{\"question\":\"How are train-time observations connected to downstream classification performance?\",\"answer\":\"The paper uses training-time signals (including changes reflected in embeddings and loss behavior) to look for correlations with classification task performance. It reports tentative correlations and notes that replication is needed.\"}]","Domain Fine-Tuning FinBERT on Finnish Histopathological Reports - Train-Time Signals and Downstream Correlations | PDF"]