[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119032-en":3,"doc-seo-119032-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119032,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","MiraBest: a data set of morphologically classified radio galaxies for machine learning","Current and next-generation astronomical surveys generate large-scale data that increasingly rely on automated machine learning. Yet standardized data sets for benchmarking machine learning performance in astronomy remain limited. MiraBest addresses this gap by providing a public, batched set of 1256 radio-loud AGN from NVSS and FIRST, filtered to 0.03 \u003C z \u003C 0.1 and manually labelled using the Fanaroff–Riley morphological scheme. The paper details construction principles, sample selection, pre-processing, dataset structure, and compares MiraBest with related literature sets. It also cross-matches to create an extended dataset of 2100 sources for broader machine-learning use.","RASTAI 2, 293–306 (2023) [https://doi.org/10.1093/rasti/rzad017](https://doi.org/10.1093/rasti/rzad017)  \nAdvance Access publication 2023 June 19  \nMiraBest: a data set of morphologically classiﬁed radio galaxies for machine learning  \nFiona A. M. Porter 1‹ and Anna M. M. Scaife 1,2  \n1Jodrell Bank Centrefor Astrophysics, Department of Physics & Astronomy, University of Manchester, Oxford Road, Manchester M13 9PL, UK  \n2 The Alan Turing Institute, Euston Road, London NW1 2DB, UK  \nAccepted 2023 May 24. Received 2023 May 18; in original form 2022 October 26  \nABSTRACT  \nThe volume of data from current and future observatories has motivated the increased development and application of automated machine learning methodologies for astronomy. However, less attention has been given to the production of standardized data sets for assessing the performance of different machine learning algorithms within astronomy and astrophysics. Here we describe in detail the MiraBest data set, a publicly available batched data set of 1256 radio-loud AGN from NVSS and FIRST, ﬁltered to 0.03 \u003C z \u003C 0.1, manually labelled by Miraghaei and Best according to the Fanaroff–Riley morphological classiﬁcation, created for machine learning applications and compatible for use with standard deep learning libraries. We outline the principles underlying the construction of the data set, the sample selection and pre-processing methodology, data set structure and composition, as well as a comparison of MiraBest to other data sets used in the literature. Existing applications that utilize the MiraBest data set are reviewed, and an extended data set of 2100 sources is created by cross-matching MiraBest with other catalogues of radio-loud AGN that have been used more widely in the literature for machine learning applications.  \nKey words: Machine Learning–astronomical data bases–radio continuum: galaxies.  \n1 INTRODUCTION  \nIn radio astronomy, morphological classiﬁcation using convolutional neural networks (CNNs) and deep learning is becoming increasingly common for object classiﬁcation, in particular with respect to the classiﬁcation of radio galaxies (see e.g. Aniyan & Thorat 2017; Alger et al. 2018; Lukic et al. 2018, 2019; Wu et al. 2018; Tang et al. 2019; Becker et al. 2021; Bowles et al. 2021; Ntwaetsile & Geach 2021; Sadeghi et al. 2021; Scaife & Porter 2021; Wang et al. 2021; Mohan et al. 2022; Slijepcevic et al. 2022, etc.). Many of these works have focused on the morphological classiﬁcation of radio galaxies following the Fanaroff–Riley classiﬁcation scheme (FR; Fanaroff & Riley 1974), used to group radio-loud active galactic nuclei (AGNs) by examining the locations of their regions of greatest luminosity relative to overall source extent. The initial scheme posited that there were two major populations of such sources–those which were corebrightened, with their peak luminosity concentrated at a radius of less than half than the overall angular size of the source from its centre (FR Type I), and those which were edge-brightened, with their peak luminosity concentrated at a radius of more than half the angular size of the source (FR Type II), and that there was a division in luminosity between the two populations at approximately 1025 Watts Hz−1 sr−1 , with edge-brightened sources having a higher intrinsic luminosity than core-brightened sources. As described, this taxonomy requires that an AGN is associated with well-resolved extended emission external to the AGN core in order tobe classifed as either FRI or FRII.  \n􀀂 E-mail: ﬁ[onamayporter@gmail.com](onamayporter@gmail.com)  \nWhile the Fanaroff–Riley scheme was initially viewed as having a very straightforward luminosity boundary between morphological classes (Fanaroff & Riley 1974), further study has shown that this is not the case, see e.g. Hardcastle & Croston (2020) for a review. In recent studies, sources have been detected which have raised questions about the use of this boundary; for example, Mingo e","cbCainfnbXaTGTpS","https://ap.wps.com/l/cbCainfnbXaTGTpS","pdf",889842,1,14,"English","en",105,"# Introduction\n## Fanaroff–Riley scheme and its limitations\n## Motivation for standardized FR training data\n## Radio surveys enabling deeper samples","[{\"question\":\"What problem does the MiraBest dataset address for machine learning in astronomy?\",\"answer\":\"It provides a standardized, publicly available dataset to benchmark and compare machine learning algorithms for astronomical morphological classification, which has been comparatively underdeveloped.\"},{\"question\":\"How is MiraBest constructed and what sources does it include?\",\"answer\":\"MiraBest is a batched dataset of 1256 radio-loud AGN drawn from NVSS and FIRST, filtered to 0.03 \\u003c z \\u003c 0.1 and manually labelled according to the Fanaroff–Riley morphological classification.\"},{\"question\":\"What extended dataset is created beyond the original MiraBest sample?\",\"answer\":\"An extended dataset of 2100 sources is produced by cross-matching MiraBest with other more widely used radio-loud AGN catalogues used for machine learning applications.\"}]","MiraBest: a data set of morphologically classified radio galaxies for machine learning | PDF",1785722013,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mirabest-a-data-set-of-morphologically-classified-radio-galaxies-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mirabest-a-data-set-of-morphologically-classified-radio-galaxies-for-machine-learning/119032/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the MiraBest dataset address for machine learning in astronomy?","Question",{"text":75,"@type":76},"It provides a standardized, publicly available dataset to benchmark and compare machine learning algorithms for astronomical morphological classification, which has been comparatively underdeveloped.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is MiraBest constructed and what sources does it include?",{"text":80,"@type":76},"MiraBest is a batched dataset of 1256 radio-loud AGN drawn from NVSS and FIRST, filtered to 0.03 \u003C z \u003C 0.1 and manually labelled according to the Fanaroff–Riley morphological classification.",{"name":82,"@type":73,"acceptedAnswer":83},"What extended dataset is created beyond the original MiraBest sample?",{"text":84,"@type":76},"An extended dataset of 2100 sources is produced by cross-matching MiraBest with other more widely used radio-loud AGN catalogues used for machine learning applications.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]