[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128418-en":3,"doc-seo-128418-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128418,8796095027276,"Valentina","https://avatar.qwps.com/avatar/d3BzX2FwX3Rlc3RfMjUxMTI2XzAxODA=",8,"Research & Report","Estimating Exoplanet Mass Using Machine Learning on Incomplete Datasets","The exoplanet archive provides extensive information on discovered extrasolar planets, yet missing values restrict statistical analysis. Planet mass is especially difficult because more than 70% of planets lack measured mass. Five machine-learning methods are compared for imputing missing properties from multidimensional incomplete datasets, focusing on planet-mass recovery and the ability to predict mass using any available subset of six or eight planet properties. Results improve even with additional incomplete data, with kNN×KDE favored for returning probability distributions enabling confidence assessment and demographic insights across different discovery methods. Open-source code accompanies the study.","arXiv :2410 .06922v1 [ astro-ph .EP] 9 Oct 2024  \nVersion October 10, 2024  \nPreprint typeset using LATEX style openjournal v. 09/06/15  \nESTIMATING EXOPLANET MASS USING MACHINE LEARNING ON INCOMPLETE DATASETS  \nFlorian Lalande 1 , ∗ , Elizabeth Tasker2 , and Kenji Doya 1  \n1 Okinawa Institute of Science and Technology. 1919-1 Tancha, Onna, Kunigami, Okinawa 904-0495, Japan and  \n2 Institute of Space and Astronautical Science, JAXA, Yoshinodai 3-1-1, Sagamihara, Kanagawa 252-5210, Japan  \nVersion October 10, 2024  \nABSTRACT  \nThe exoplanet archive is an incredible resource of information on the properties of discovered extrasolar planets, but statistical analysis has been limited by the number of missing values. One of the most informative bulk properties is planet mass, which is particularly challenging to measure with more than 70% of discovered planets with no measured value. We compare the capabilities of five different machine learning algorithms that can utilize multidimensional incomplete datasets to estimate missing properties for imputing planet mass. The results are compared when using a partial subset of the archive with a complete set of six planet properties, and where all planet discoveries are leveraged in an incomplete set of six and eight planet properties. We find that imputation results improve with more data even when the additional data is incomplete, and allows a mass prediction for any planet regardless of which properties are known. Our favored algorithm is the newly developed kNN×KDE, which can return a probability distribution for the imputed properties. The shape of this distribution can indicate the algorithm’s level of confidence, and also inform on the underlying demographics of the exoplanet population. We demonstrate how the distributions can be interpreted with a series of examples for planets where the discovery was made with either the transit method, or radial velocity method. Finally, we test the generative capability of the kNN×KDE to create a large synthetic population of planets based on the archive, and identify potential categories of planets from groups of properties in the multidimensional space. All codes are Open Sourcea.  \nSubject headings: Exoplanet catalogs, Astronomy databases, Astrostatistics tools, Computational methods  \n1. INTRODUCTION  \nSince the first discoveries in the early 1990s, over 5,500 planets have been discovered outside our Solar System (Wolszczan & Frail 1992 ; Mayor & Queloz 1995 ; Akeson et al. 2013) . While the planets orbiting our Sun can be categorized as either rocky or gaseous simply depending on their orbital period, the myriad of sizes and orbits of the planets detected around other stars point toa multitude of formation pathways for planets that are influenced by a wide range of environmental factors.  \nDedicated survey missions such as Convection, Rotation et Transits planétaires (CoRoT), the Kepler space telescope, and the Transiting Exoplanet Survey Satellite (TESS), alongside ground-based search instruments and programs that include the High Accuracy Radial Velocity Planet Searcher (HARPS), Wide Angle Search for Planets (WASP) and Optical Gravitational Lensing Experiment (OGLE) are trying to build a census of planet types. This has resulted in the construction of a large archive of data for the properties of the discovered planets. This exoplanet archive is an invaluable resource for identifying patterns and trends between planet and stellar properties. Such relationships can be used to estimate properties of planets that have not (and often cannot) be measured, allowing a more complete picture of planet diversity and information to select the most promising targets for time-consuming atmospheric characterization studies by instruments such as the James Webb Space Telescope. However, making full use of the archive has turned out to be challenging.  \nOne of the principal difficulties is that the archive con-  \n∗ [E-mail:](E-mail: florian.lalande@oi","cbCaigfnJVmkdmGo","https://ap.wps.com/l/cbCaigfnJVmkdmGo","pdf",5991006,2,1,30,"English","en",105,"# Abstract\n# Introduction\n## Exoplanet archive and missing values\n## Discovery techniques and sparse observations","[{\"question\":\"Why is estimating exoplanet mass difficult in the exoplanet archive?\",\"answer\":\"Most discovered planets have no measured mass, creating a dataset with over 70% missing mass values. Different discovery techniques also record only small subsets of properties, making joint statistical analysis challenging.\"},{\"question\":\"What is the main goal of comparing five machine learning algorithms?\",\"answer\":\"To evaluate how effectively each algorithm can impute missing planet properties—especially mass—using multidimensional incomplete datasets, including cases where only partial subsets of properties are available.\"},{\"question\":\"How does the favored kNN×KDE approach help interpret the results?\",\"answer\":\"kNN×KDE returns a probability distribution for imputed properties. The distribution’s shape reflects the model’s confidence and can also reveal demographic structure in the exoplanet population.\"}]","Estimating Exoplanet Mass Using Machine Learning on Incomplete Datasets | PDF",1785947400,76,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"estimating-exoplanet-mass-using-machine-learning-on-incomplete-datasets","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/estimating-exoplanet-mass-using-machine-learning-on-incomplete-datasets/128418/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is estimating exoplanet mass difficult in the exoplanet archive?","Question",{"text":76,"@type":77},"Most discovered planets have no measured mass, creating a dataset with over 70% missing mass values. Different discovery techniques also record only small subsets of properties, making joint statistical analysis challenging.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the main goal of comparing five machine learning algorithms?",{"text":81,"@type":77},"To evaluate how effectively each algorithm can impute missing planet properties—especially mass—using multidimensional incomplete datasets, including cases where only partial subsets of properties are available.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the favored kNN×KDE approach help interpret the results?",{"text":85,"@type":77},"kNN×KDE returns a probability distribution for imputed properties. The distribution’s shape reflects the model’s confidence and can also reveal demographic structure in the exoplanet population.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":22,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]