[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119378-en":3,"doc-seo-119378-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119378,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","High-Throughput Phenotyping using Computer Vision and Machine Learning - SMC Data Challenge #1","High-throughput phenotyping enables non-destructive, efficient evaluation of plant phenotypes and has increasingly been enhanced with machine learning to process large image datasets and extract targeted traits. This study uses a dataset of 1,672 Populus trichocarpa images containing physical white labels for treatment, block, row, position, and genotype, along with EXIF-encoded metadata. OCR reads label text for spreadsheet-ready organization, segmentation and classifiers predict morphological traits, and a treatment prediction model is trained from the classifications. EXIF tags are analyzed for leaf size relationships, while missing EXIF information limits certain assessments, suggesting directions for future improvements.","10 Jul 2024  \nHigh-Throughput Phenotyping using Computer Vision and Machine Learning  \nSMC Data Challenge \\#1  \nVivaan Singhvia , Langalibalele Lungaa , Pragya Nidhia , Chris Keuma , Varrun Prakasha  \naFarragut High School, 11237 Kingston Pike, Knoxville, 37922, Tennessee, United States  \nAbstract  \nHigh-throughput phenotyping refers to the non-destructive and efficient evaluation of plant phenotypes. In recent years, it has been coupled with machine learning in order to improve the process of phenotyping plants by increasing efficiency in handling large datasets and developing methods for the extraction of specific traits. Previous studies have developed methods to advance these challenges through the application of deep neural networks in tandem with automated cameras; however, the datasets being studied often excluded physical labels. In this study, we used a dataset provided by Oak Ridge National Laboratory with 1,672 images of Populus Trichocarpa with white labels displaying treatment (control or drought), block, row, position, and genotype. Optical character recognition (OCR) was used to read these labels on the plants, image segmentation techniques in conjunction with machine learning algorithms were used for morphological classifications, machine learning models were used to predict treatment based on those classifications, and analyzed encoded EXIF tags were used for the purpose of finding leaf size and correlations between phenotypes. We found that our OCR model had an accuracy of 94.31% for non-null text extractions, allowing for the information to be accurately placed in a spreadsheet. Our classification models identified leaf shape, color, and level of brown splotches with an average accuracy of 62.82%, and plant treatment with an accuracy of 60.08% . Finally, we identified a few crucial  \npieces of information absent from the EXIF tags that prevented the assessment of the leaf size. There was also missing information that prevented the assessment of correlations between phenotypes and conditions. However, future studies could improve upon this to allow for the assessment of these features. The use of machine learning and computer vision in high-throughput phenotyping has shown to be effective in analyzing large plant datasets, leading to a more comprehensive phenotype analysis in plants and showing potential in various agricultural and environmental applications.  \nKeywords: computer vision, machine learning, high-throughput phenotyping  \n1. Introduction  \n1.1. Background Information  \nHigh-throughput phenotyping is defined as a breakthrough technology used in plant biology and agriculture to examine and assess plants’”anatomical, ontological, physiological, and biochemical features” through the use of images. Previous studies have shown its potential as a noninvasive replacement for traditional on-field techniques used to extract important phenotypic data. In recent years, this potential has been further intensified by coupling image-based phenotyping techniques with machine learning algorithms. This new approach has enabled the extraction of phenotypes from ”complex” plant image datasets that were previously challenging to analyze efficiently. Furthermore, newer studies have shown that the technology has potential to identifying correlations between phenotype, genotype, and environmental metadata.  \ncontains physical labels within the images, each containing important information on treatment (control or drought), block, row, position, and genotype. As a result, this study will address a new aspect of image-based phenotyping, namely, extracting data from the white labels to identify correlations between leaves’ phenotypes and data embedded within the white labels. During this project, we aim to answer the following questions:  \n1. Is it possible to use optical character recognition (OCR) or machine learning techniques to “read” the label on each tag and generate a spreadsheet containing the treatment, block, ro","cbCaiuhkcGJ9qv2U","https://ap.wps.com/l/cbCaiuhkcGJ9qv2U","pdf",6921670,1,11,"English","en",105,"# Abstract\n# 1. Introduction\n## 1.1. Background Information\n## 1.2. Research Objective\n# 2. Related Works\n## 2.1. Advancements in High Throughput Phenotyping Techniques","[{\"question\":\"What problem does this study address in high-throughput phenotyping?\",\"answer\":\"It targets efficient phenotyping from plant images while addressing limitations of datasets that often lack physical labels needed to extract treatment and other structured metadata.\"},{\"question\":\"How is the physical label information extracted and used?\",\"answer\":\"Optical character recognition (OCR) reads the white labels on plants, and the extracted information is organized into a spreadsheet containing treatment, block, row, position, and genotype.\"},{\"question\":\"Which traits and outcomes were predicted, and what were the reported performance levels?\",\"answer\":\"The models identify leaf shape, color, and brown spot levels (average accuracy 62.82%) and predict plant treatment based on those classifications (average accuracy 60.08%), while OCR achieved 94.31% accuracy for non-null text extraction.\"}]","High-Throughput Phenotyping using Computer Vision and Machine Learning - SMC Data Challenge #1 | PDF",1785724009,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"high-throughput-phenotyping-using-computer-vision-and-machine-learning-smc-data-challenge-1","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/high-throughput-phenotyping-using-computer-vision-and-machine-learning-smc-data-challenge-1/119378/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this study address in high-throughput phenotyping?","Question",{"text":75,"@type":76},"It targets efficient phenotyping from plant images while addressing limitations of datasets that often lack physical labels needed to extract treatment and other structured metadata.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the physical label information extracted and used?",{"text":80,"@type":76},"Optical character recognition (OCR) reads the white labels on plants, and the extracted information is organized into a spreadsheet containing treatment, block, row, position, and genotype.",{"name":82,"@type":73,"acceptedAnswer":83},"Which traits and outcomes were predicted, and what were the reported performance levels?",{"text":84,"@type":76},"The models identify leaf shape, color, and brown spot levels (average accuracy 62.82%) and predict plant treatment based on those classifications (average accuracy 60.08%), while OCR achieved 94.31% accuracy for non-null text extraction.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]