[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-seo-307654-105":53,"doc-detail-307654-en":126},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":54,"data":55},"ok",{"site_id":56,"language":57,"slug":58,"title":59,"keywords":60,"description":61,"schema_data":62,"social_meta":119,"head_meta":121,"extra_data":123,"updated_unix":125},105,"en","statllama-multi-stage-training-for-domain-optimized-statistical-large-language-models-abstract","StatLLaMA - Multi-Stage Training for Domain-Optimized Statistical Large Language Models - Abstract","","This study investigates how to efficiently build a domain-specialized large language model for statistics using the lightweight LLaMA-3.2-3B family as the foundation model. It systematically compares three multi-stage training pipelines across continual pretraining, supervised fine-tuning, and RLHF preference alignment. Pipelines starting from a base model without instruction-following fail to develop meaningful statistical reasoning. In contrast, starting from LLaMA-3.2-3B-Instruct enables effective domain specialization with clear trade-offs. Direct preference optimization yields stable RLHF alignment, and StatLLaMA attains strong balanced performance on reasoning and statistical benchmarks.",{"@graph":63,"@context":118},[64,80,101],{"@type":65,"itemListElement":66},"BreadcrumbList",[67,71,74,77],{"item":68,"name":69,"@type":70,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":72,"name":10,"@type":70,"position":73},"https://docshare.wps.com/template/",2,{"item":75,"name":51,"@type":70,"position":76},"https://docshare.wps.com/template/general/",3,{"item":78,"name":59,"@type":70,"position":79},"https://docshare.wps.com/template/statllama-multi-stage-training-for-domain-optimized-statistical-large-language-models-abstract/307654/",4,{"url":78,"name":59,"@type":81,"image":82,"author":87,"headline":59,"publisher":90,"fileFormat":93,"inLanguage":57,"description":61,"dateModified":94,"datePublished":95,"encodingFormat":93,"isAccessibleForFree":96,"interactionStatistic":97},"DigitalDocument",{"url":83,"@type":84,"width":85,"height":86},"https://docshare.wps.com/thumbnails/statllama-multi-stage-training-for-domain-optimized-statistical-large-language-models-abstract/307654.png","ImageObject",442,249,{"name":88,"@type":89},"Lucas Martin","Person",{"url":68,"name":91,"@type":92},"DocShare","Organization","application/pdf","2026-09-20","2026-09-19",true,{"@type":98,"interactionType":99,"userInteractionCount":73},"InteractionCounter",{"@type":100},"ViewAction",{"@type":102,"mainEntity":103},"FAQPage",[104,110,114],{"name":105,"@type":106,"acceptedAnswer":107},"What is the main goal of StatLLaMA?","Question",{"text":108,"@type":109},"To build an efficient statistical large language model by leveraging multi-stage training that improves domain expertise while preserving general reasoning and language capabilities.","Answer",{"name":111,"@type":106,"acceptedAnswer":112},"How do the three training pipelines differ in their effectiveness?",{"text":113,"@type":109},"Pipelines that start from a base foundation model without instruction-following fail to develop meaningful statistical reasoning, while starting from LLaMA-3.2-3B-Instruct enables effective domain specialization with measurable trade-offs.",{"name":115,"@type":106,"acceptedAnswer":116},"Why is preference alignment included, and what methods are compared?",{"text":117,"@type":109},"Preference alignment addresses challenges in matching human judgments for nuanced reasoning. The document discusses RLHF and also compares more efficient alternatives such as direct preference optimization (DPO) and group relative policy optimization (GRPO).","https://schema.org",{"og:url":78,"og:type":120,"og:title":59,"og:site_name":91,"og:description":61},"article",{"robots":122,"canonical":78},"index,follow",{"doc_id":124,"site_id":56},307654,1789852602,{"code":4,"msg":5,"data":127},{"doc_id":124,"user_id":128,"nickname":88,"user_avatar":129,"doc_module":9,"category_id":50,"category_name":51,"doc_title":59,"doc_description":61,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":73,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":135,"language":136,"language_code":57,"site_id":56,"html_lang":57,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":61,"update_tm":125,"read_time":20},8796095360427,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","June 2026, Volume VI, Issue IV. doi: 10.52933/jdssv.v6i4.171  \nStatLLaMA: Multi-Stage Training for Domain-Optimized Statistical Large Language Models  \nJing-Yi Zeng  \nNational Yang Ming Chiao Tung University  \nGuan-Hua Huang  \nNational Yang Ming Chiao Tung University  \nAbstract  \nThis study investigates how to efficiently build a domain-specialized large language model (LLM) for statistics using the lightweight LLaMA-3.2-3B family as the foundation model (FM) . We systematically compare three multi-stage training pipelines—starting from a base FM with no instruction-following capability, a base FM augmented with post-hoc instruction tuning, and an instructiontuned FM with strong general reasoning abilities—across continual pretraining, supervised fine-tuning (SFT), and reinforcement learning from human feedback (RLHF) preference alignment. Results show that pipelines beginning with abase FM fail to develop meaningful statistical reasoning, even after extensive instruction tuning, SFT, or RLHF alignment. In contrast, starting from LLaMA- 3.2-3B-Instruct enables effective domain specialization. A comprehensive evaluation of SFT variants reveals clear trade-offs between domain expertise and general reasoning ability. We further demonstrate that direct preference optimization provides stable and effective RLHF preference alignment. The final model, StatLLaMA, achieves strong and balanced performance on benchmarks of mathematical reasoning, common-sense reasoning, and statistical expertise, offering a practical blueprint for developing resource-efficient statistical LLMs. The code is available at [https://github.com/HuangDLab/StatLLaMA](https://github.com/HuangDLab/StatLLaMA).  \n2 StatLLaMA, Version 1  \nKeywords: Foundation model, supervised fine-tuning, continual pretraining, instruction tuning, reinforcement learning from human feedback.  \n1 . Introduction  \nLarge language models (LLMs) based on Transformer architectures have become a central tool in modern natural language processing. Models such as BERT (Devlin et al. 2019), GPT (Radford et al. 2018 , 2019 ; Brown et al. 2020), and the LLaMA family (Touvron et al. 2023a,b; Grattafiori et al. 2024) have demonstrated strong performance across a wide range of general-purpose language tasks, including reasoning, summarization, and question answering. Trained on massive text corpora, these foundation models exhibit broad linguistic competence and extensive general knowledge. However, their effectiveness in specialized technical domains—particularly statistics—remains limited.  \nStatistical reasoning requires precise use of domain-specific terminology, adherence to formal definitions, and careful multi-step analytical reasoning. In practice, generalpurpose LLMs often produce explanations that are shallow, imprecise, or conceptually incorrect when applied to statistical tasks. This gap raises an important applied question: how can statistical knowledge and reasoning be systematically integrated into LLMs without degrading their general language and reasoning abilities? This study addresses this question by developing and evaluating a resource-efficient training framework for constructing a statistical LLM. We focus on the lightweight LLaMA-3 .2-3B model, motivated by the practical need for deployable models under constrained computational resources. Our objective is to build a model that achieves strong performance on statistical reasoning tasks while maintaining robust general-purpose capabilities.  \nSupervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream tasks. However, full fine-tuning becomes prohibitively expensive for large models and is inefficient when multiple domain-specific variants are required. Parameter-efficient fine-tuning (PEFT) methods, particularly low-rank adaptation (LoRA) (Hu et al. 2021), address this limitation by updating only a small subset of parameters while keeping the base model fixed. These techniques s","cbCaikQmVdQoIoT4","https://ap.wps.com/l/cbCaikQmVdQoIoT4","pdf",1717358,39,"English","# Abstract\n# Introduction\n## Motivation and background on LLMs\n## Supervised fine-tuning and parameter-efficient approaches\n## Continual pretraining vs instruction tuning\n## Preference alignment: RLHF, DPO, and GRPO","[{\"question\":\"What is the main goal of StatLLaMA?\",\"answer\":\"To build an efficient statistical large language model by leveraging multi-stage training that improves domain expertise while preserving general reasoning and language capabilities.\"},{\"question\":\"How do the three training pipelines differ in their effectiveness?\",\"answer\":\"Pipelines that start from a base foundation model without instruction-following fail to develop meaningful statistical reasoning, while starting from LLaMA-3.2-3B-Instruct enables effective domain specialization with measurable trade-offs.\"},{\"question\":\"Why is preference alignment included, and what methods are compared?\",\"answer\":\"Preference alignment addresses challenges in matching human judgments for nuanced reasoning. The document discusses RLHF and also compares more efficient alternatives such as direct preference optimization (DPO) and group relative policy optimization (GRPO).\"}]","StatLLaMA - Multi-Stage Training for Domain-Optimized Statistical Large Language Models - Abstract | PDF"]