[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127616-en":3,"doc-seo-127616-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127616,549768064778,"Finn","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",6,"Technology","Machine Learning Techniques for Fish Breeding - Decision Making","New Zealand breeding programs for Australasian snapper focus on selecting individuals that produce quicker-maturing, high-quality offspring. Genomic datasets support this goal but contain missing values across many features, requiring careful data imputation to preserve gene signal for growth-rate prediction. The study evaluates how five imputation strategies affect classification using a Random Forest model, comparing standard and domain-based KNN and MICE variants with varied parameter settings. Results show all methods remain robust and yield similar accuracies, while domain-based imputation shows a modest advantage confirmed by overall significance testing.","Machine Learning Techniques for Fish Breeding  \nDecision Making  \nRose Taylor School of Engineering and Computer  \nScience Victoria University of Wellington  \nWellington, New Zealand  \n[taylorrose3@myvuw.ac.nz](taylorrose3@myvuw.ac.nz)  \nAbstract— The New Zealand Institute for Plant and Food Research has been working on creating breeding programs for the Australasian Snapper (Chrysophrys auratus) to breed snappers that mature faster and are high quality. One of the breeding program goals is to select individuals that produce quick-to-mature offspring. To accomplish this, they collected the genomic makeup of snappers into a dataset. However, the collected data has missing values in some features, which require imputation to enable use of those features to classify fish that grow faster and slower. As the genes responsible for controlling the growth rate in Snapper are currently unknown, the dataset must maintain as many of the features as possible to enable identification of the genes most likely to control the snappers' growth rate. This project investigated whether the data imputation methods used impacted the ability of a machine learning classifier to predict the growth rate and, if so, how different imputation methods performed. This project implemented five imputation methods, specifically Most Frequent imputation, K-Nearest Neighbour (KNN) imputation, Multiple Imputation by Chained Equations (MICE), a KNN approach using domain information, and a cascading KNN imputation method using domain information. The KNN and MICE approaches have two different parameter settings for imputation. This project evaluated these imputation techniques using a Random Forest classifier. The results showed that all imputation methods are robust to the test train split and random state used in the random forest classifier. The classification accuracies were similar between the imputation methods. Despite differences being displayed in split datasets, the complete datasets p-value calculations confirmed no significant differences in overall result. These results indicated that domain-based imputation approaches did perform better than other imputation techniques indicating that using domain-based imputation techniques could improve the overall classification accuracy. Lack of significant differences between the classification accuracies are caused by the number of features being so great that there is little overlap in the features selected by the Random Forest classifier and the features that are selected by the majority of the trees help account for majority of the classification accuracy.  \nKeywords—data imputation, machine learning, fish, KNN, MICE  \nI. INTRODUCTION  \nThis project is part of the Australasian Snapper (Chrysophrys auratus) breeding program begun in 2016 by The New Zealand Institute for Plant and Food Research Limited [1] . The breeding program aims to identify the genotypes responsible for controlling the growth rate in snapper fish so that when planners breed snapper, the offspring grow to a harvestable size faster [1] . The snapper growth rate is slow. It can take 3-5 years for a snapper to grow large enough to be legally caught commercially for food [2],  \n[3] . This project aims to solve the problem of how to handle missing data within the genetic dataset. Since the genes responsible for influencing the growth rate in snapper are, as yet, unknown, the handling of missing data is vital to creating a model that can accurately predict the growth rate of a snapper based on its DNA [1] .  \nThis project investigated whether the use of different data imputation methods impacts the ability of a machine learning classifier to predict the growth rate and, if so, how different imputation methods perform. The project implemented and evaluated different techniques for imputing the missing data values within the original dataset. This project is part of an existing project being worked on by multiple researchers, both within Victoria U","cbCaijMGdIc7sXEG","https://ap.wps.com/l/cbCaijMGdIc7sXEG","pdf",410270,2,1,12,"English","en",105,"# Introduction\n## Missing data in snapper genomic datasets\n## Goal and prior work\n## Motivation and evaluation approach","[{\"question\":\"Why is data imputation required in the snapper genomic dataset?\",\"answer\":\"Missing values exist in genomic features, and features cannot be removed because the growth-rate genes are still unknown. Imputation enables classifiers to use the full feature set for predicting faster vs slower growth.\"},{\"question\":\"Which imputation methods were implemented in the project?\",\"answer\":\"The study implemented Most Frequent imputation, K-Nearest Neighbour (KNN) imputation, Multiple Imputation by Chained Equations (MICE), a domain-informed KNN approach, and a cascading domain-informed KNN imputation method.\"},{\"question\":\"How were imputation methods evaluated for predicting growth rate?\",\"answer\":\"A Random Forest classifier was used, and performance was assessed across different train/test split settings and random states. Overall results were checked with p-value calculations to confirm whether differences between imputation methods were significant.\"}]","Machine Learning Techniques for Fish Breeding - Decision Making | PDF",1785940294,30,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"machine-learning-techniques-for-fish-breeding-decision-making","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/machine-learning-techniques-for-fish-breeding-decision-making/127616/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is data imputation required in the snapper genomic dataset?","Question",{"text":76,"@type":77},"Missing values exist in genomic features, and features cannot be removed because the growth-rate genes are still unknown. Imputation enables classifiers to use the full feature set for predicting faster vs slower growth.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which imputation methods were implemented in the project?",{"text":81,"@type":77},"The study implemented Most Frequent imputation, K-Nearest Neighbour (KNN) imputation, Multiple Imputation by Chained Equations (MICE), a domain-informed KNN approach, and a cascading domain-informed KNN imputation method.",{"name":83,"@type":74,"acceptedAnswer":84},"How were imputation methods evaluated for predicting growth rate?",{"text":85,"@type":77},"A Random Forest classifier was used, and performance was assessed across different train/test split settings and random states. Overall results were checked with p-value calculations to confirm whether differences between imputation methods were significant.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,114,119,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":112,"slug":113},50,"technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":30,"slug":122},8,"Research & Report","research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]