[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127583-en":3,"doc-seo-127583-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127583,549768064622,"Anda","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Machine learning of kinetic energy densities with target and feature averaging - better results with fewer training data","Machine learning of kinetic energy functionals, especially kinetic energy density (KED) functionals, enables constructing kinetic energy functionals for orbital-free density functional theory. Neural networks and kernel methods such as Gaussian process regression learn Kohn-Sham KED from density-based descriptors, but uneven and highly nonuniform datasets hinder training, increasing the need for samples and risking overfitting. Partially averaged density-dependent variables and KED are used to smooth sampling difficulty while preserving spatial structure. Gaussian process regression on partially spatially averaged terms of the 4th order gradient expansion and the Kohn-Sham effective potential yields accurate, stable kinetic energy models for Al, Mg, and Si from as few as 2000 samples, reaching about 1% accuracy in energy–volume dependence simultaneously.","Machine learning of kinetic energy densities with target and feature averaging: better results with fewer training data  \nSergei Manzhos 1,a, Johann Lüder2,3,4, Manabu Ihara 1  \n1 School of Materials and Chemical Technology, Tokyo Institute of Technology, Ookayama 2-12-1, Meguro-ku, Tokyo 152-8552 Japan  \n2 Department of Materials and Optoelectronic Science, National Sun Yat-sen University, 80424, No. 70, Lien-Hai Road, Kaohsiung, Taiwan R.O.C.  \n3 Center of Crystal Research, National Sun Yat-sen University, 80424, No. 70, Lien-Hai Road, Kaohsiung, Taiwan R.O.C.  \n4 Center for Theoretical and Computational Physics, National Sun Yat-Sen University, Kaohsiung 80424, Taiwan  \nAbstract  \nMachine learning of kinetic energy functionals (KEF), in particular kinetic energy density (KED) functionals, has recently attracted attention as a promising way to construct KEFs for orbital-free density functional theory (OF-DFT) . Neural networks (NN) and kernel methods including Gaussian process regression (GPR) have been used to learn Kohn-Sham (KS) KED from density-based descriptors derived from KSDFT calculations. The descriptors are typically expressed as functions of different powers and derivatives ofthe electron density. This can generate large and extremely unevenly distributed datasets, which complicates effective application of machine learning techniques. Very uneven data distributions require many training data points, can cause overfitting, and ultimately lower the quality of a ML KED model. We show  \na Author to whom correspondence should be [addressed. Email: Manzhos.s.aa@m.titech.ac.jp](addressed. Email: Manzhos.s.aa@m.titech.ac.jp)  \nthat one can produce more accurate ML models from fewer data by working with partially averaged density-dependent variables and KED. Averaging palliates the issue of very uneven data distributions and associated difficulties of sampling, while retaining enough spatial structure necessary for working within the paradigm ofKEDF. We use GPR as a function of partially spatially averaged terms of the 4th order gradient expansion and the Kohn-Sham effective potential and obtain accurate and stable (with respect to different random choices of training points) kinetic energy models for Al, Mg, and Si simultaneously from as few as 2000 samples (about 0.3% of the total KSDFT data) . In particular, accuracies on the order of 1% in a measure of the quality of energy-volume dependence 􀜤 ′ = 􀮾~~ ~~(􀯏0~~ ~~−Δ􀯏)~~ ~~−(2Δ􀮾􀯏⁄(􀯏􀯏00))+2􀮾~~ ~~(􀯏0~~ ~~+Δ􀯏) are obtained simultaneously for all three materials.  \n1 Introduction  \nOrbital-free density functional theory (OF-DFT)1–3 has the potential to revolutionize the field of computational materials modeling by making large-scale DFT calculations fast and routine. In Kohn-Sham (KS) DFT4,5 currently dominating ab initio materials modeling, the energy of a system of 􀜰􀯘􀯟 electrons (we neglect spin without loss of generality) is computed as a functional of the electron density 􀟩(􀢘),  \n􀯇 􀳐􀳗  \n􀜧 = − 12 ∑ ∫ 􀟰 (􀢘)Δ􀟰􀯜 (􀢘)􀝀􀢘 + ∫ 􀜸􀯜􀯢􀯡 (􀢘)􀟩 (􀢘)􀝀􀢘 + 12 ∬ 􀟩~~ ~~|(􀢘􀢘􀢘(􀢘′|′) 􀝀􀢘􀝀􀢘′ + 􀜧􀯑􀮼 [􀟩 (􀢘)]􀯜 = 1  \n(1.1)  \nwhere the orbitals 􀟰􀯜 (􀢘) are the solutions ofthe Kohn-Sham equation  \n1  \n− Δ􀟰􀯜 (􀢘) + 􀜸􀯘􀯙􀯙 [􀟩 (􀢘)]􀟰 􀯜 (􀢘) = 􀟳 􀯜 􀟰 􀯜 (􀢘)  \n2  \n􀜸􀯘􀯙􀯙 = 􀜸􀯜􀯢􀯡 (􀢘) + ∫ ~~ ~~|􀟩􀢘~~ ~~′􀢘)′~~ ~~| 􀝀􀢘′ + 􀜸􀯑􀮼 [􀟩 (􀢘)]  \n(1.2)  \nWe use atomic units unless stated otherwise. Here 􀜸􀯜􀯢􀯡 (􀢘) is the potential due to atomic nuclei, 􀜧􀯑􀮼 [􀟩 (􀢘)] is the exchange-correlation energy, and 􀜸􀯑􀮼 [􀟩 (􀢘)] = 􀰋􀮾􀳉􀰋􀲴􀰘[(􀰘􀢘)(􀢘)] is the exchange-correlation potential. The inclusion of 􀜸􀯑􀮼 into the effective potential 􀜸􀯘􀯙􀯙 is to make sure that ∑􀳐􀳗1|􀟰􀯜 (􀢘)|2 = 􀟩 (􀢘), the true electron density.5 The need to compute the orbitals causes a near-cubic scaling of the computational cost with system size, which is further exacerbated by the need to ensure self-consistent convergence of orbital-dependent 􀜸􀯘􀯙􀯙 [􀟩 (􀢘) = ∑􀳐􀳗1|􀟰􀯜 (􀢘)|2], as the Eq. (1 .2) is typically solved as a linear ODE for a given 􀟩 (􀢘) . As a result, routinely doable calculatio","cbCaissWPHz89Q78","https://ap.wps.com/l/cbCaissWPHz89Q78","pdf",703502,2,1,28,"English","en",105,"# Introduction\n## Orbital-free density functional theory and kinetic energy density functionals\n## Challenges from uneven training data distributions","[{\"question\":\"Why do uneven density-based descriptors complicate machine learning for kinetic energy density functionals?\",\"answer\":\"The descriptors lead to datasets that are extremely unevenly distributed, making sampling inefficient. This generally increases the number of training points required, can cause overfitting, and reduces model quality.\"},{\"question\":\"How does target and feature averaging help reduce the amount of training data needed?\",\"answer\":\"Partially averaging density-dependent variables and KED alleviates uneven data distributions and sampling difficulties. It preserves enough spatial structure to remain compatible with the kinetic energy functional (KEDF) framework.\"},{\"question\":\"What modeling method is used to learn the kinetic energy densities, and how is stability assessed?\",\"answer\":\"Gaussian process regression is used with partially spatially averaged terms of the 4th order gradient expansion and the Kohn-Sham effective potential. Stability is evaluated with respect to different random choices of training points.\"}]","Machine learning of kinetic energy densities with target and feature averaging - better results with fewer training data | PDF",1785940113,71,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"machine-learning-of-kinetic-energy-densities-with-target-and-feature-averaging-better-results-with-fewer-training-data","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/machine-learning-of-kinetic-energy-densities-with-target-and-feature-averaging-better-results-with-fewer-training-data/127583/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do uneven density-based descriptors complicate machine learning for kinetic energy density functionals?","Question",{"text":76,"@type":77},"The descriptors lead to datasets that are extremely unevenly distributed, making sampling inefficient. This generally increases the number of training points required, can cause overfitting, and reduces model quality.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does target and feature averaging help reduce the amount of training data needed?",{"text":81,"@type":77},"Partially averaging density-dependent variables and KED alleviates uneven data distributions and sampling difficulties. It preserves enough spatial structure to remain compatible with the kinetic energy functional (KEDF) framework.",{"name":83,"@type":74,"acceptedAnswer":84},"What modeling method is used to learn the kinetic energy densities, and how is stability assessed?",{"text":85,"@type":77},"Gaussian process regression is used with partially spatially averaged terms of the 4th order gradient expansion and the Kohn-Sham effective potential. Stability is evaluated with respect to different random choices of training points.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]