[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84254-en":3,"doc-seo-84254-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84254,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","An optimal control approach for neural network architecture adaptation with a posteriori error estimation","A novel method adapts neural network architecture along network depth using a posteriori error estimation. Neural network training is formulated as a continuous-time optimal control problem, yielding rigorous layer-wise approximation error estimates. The error decomposition supports a principled depth refinement strategy by inserting new layers at positions of maximum estimated error to capture nonlinear solution structure. We derive computable upper bounds via dual weighted residual methodology and provide explicit interval-wise error bounds for targeted architecture refinement. Experiments on scientific datasets, including learning the observable-to-parameter map for Navier-Stokes, show improved generalization over existing adaptation methods.","Crossmark  \nRECEIVED  \ndd Month yyyy  \nREVISED  \ndd Month yyyy  \narXiv :2607 .07637v 1 [ cs .LG] 8 Jul 2026  \nPAPER  \nAn optimal control approach for neural network architecture adaptation with a posteriori error estimation  \nC G Krishnanunni 1 , Thomas Scott 1 , Tan Bui-Thanh 1 ,∗  \n1 Department of Aerospace Engineering & Engineering Mechanics, UT Austin  \n∗ Oden Institute for Computational Engineering and Sciences, UT Austin.  \nE-mail: [krishnanunni@utexas.edu](krishnanunni@utexas.edu)  \nKeywords: Neural architecture adaptation, A posteriori error estimation, Optimal control theory.  \nAbstract  \nThis work presents a novel approach for adapting neural network architecture along the depth based on a posteriori error estimation. By formulating neural network training as a continuous-time optimal control problem, we derive rigorous error estimates that quantify how approximation error distributes across network layers. This error decomposition enables a principled depth adaptation strategy: new layers are inserted at locations of maximum estimated error, allowing the network to efficiently capture complex, nonlinear variations in the underlying problem. Our framework introduces a novel network architecture that treats weights and biases as piecewise linear functions varying across layers, with the error estimator bounding the discrepancy between this discrete representation and the true continuous optimal control solution. The approach leverages dual weighted residual methodology from finite element analysis to derive computable upper bounds on the functional error. A key theoretical contribution is the derivation of explicit error bounds that decompose the total approximation error into interval-wise contributions, providing a rigorous basis for targeted architecture refinement. We demonstrate the effectiveness of our method on scientific datasets, including learning the observable-to-parameter map for the Navier-Stokes equation. Numerical results reveal that our approach consistently outperforms existing architecture adaptation methods in terms of generalization performance.  \n1 Introduction  \nDepth of a neural network plays a key role in the observed empirical success of deep learning. Stacking layers lets a network build successively more abstract features out of raw input data [1, 2] . Several of the architectures responsible for the largest jumps in benchmark performance over the past decade differ from their predecessors mainly in how many layers they contain [3, 4] . A key question is how many layers does a given task actually require, and how should the available parameters be distributed among them? In practice this question is still settled almost entirely by trial and error ora metaheuristic optimization-based search procedure [5, 6] . A network with too little depth and small width may be unable to represent the target function at all, whereas an unnecessarily deep network wastes compute and is more prone to overfitting. A more systematic, less wasteful approach to setting network depth therefore has clear practical value.  \nTwo broad families of methods have been proposed to automate this choice. The first treats architecture selection as a search problem. A population or sequence of candidate networks is proposed, evaluated using evolutionary computation [5, 6], reinforcement learning [7], or randomized search [8], and the best candidate is retained. These neural architecture search (NAS) methods can produce excellent architectures, but doing so typically requires training large numbers of candidates which makes the overall procedure costly, particularly when a single training run is itself expensive. The second family avoids searching a discrete space of architectures altogether. Instead, a single small network is trained and then grown, with new parameters added as training proceeds rather than fixed in advance [9, 10 , 11 , 12 , 13] .  \nWithin this second family, the bulk of prior work targets towards growing t","cbCaisvsbHAXxPNn","https://ap.wps.com/l/cbCaisvsbHAXxPNn","pdf",5076669,2,1,22,"English","en",105,"# Introduction\n## Neural network depth as a key factor\n## Two families of automated architecture selection\n## Prior work on growing width vs growing depth\n## Missing quantitative link and the paper’s contribution","[{\"question\":\"How does the approach adapt neural network depth during training?\",\"answer\":\"Training is reformulated as a continuous-time optimal control problem, and posteriori error estimates determine where to insert new layers so added capacity targets the largest estimated approximation errors.\"},{\"question\":\"What is the role of dual weighted residual methodology in this framework?\",\"answer\":\"It is used to derive computable upper bounds on the functional (approximation) error, enabling practical estimation rather than only qualitative guidance.\"},{\"question\":\"What datasets and task are used to demonstrate the method?\",\"answer\":\"The method is evaluated on scientific datasets, including learning the observable-to-parameter map for the Navier-Stokes equation, where it improves generalization compared with existing architecture adaptation methods.\"}]",1784194396,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"an-optimal-control-approach-for-neural-network-architecture-adaptation-with-a-posteriori-error-estimation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/an-optimal-control-approach-for-neural-network-architecture-adaptation-with-a-posteriori-error-estimation/84254/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the approach adapt neural network depth during training?","Question",{"text":75,"@type":76},"Training is reformulated as a continuous-time optimal control problem, and posteriori error estimates determine where to insert new layers so added capacity targets the largest estimated approximation errors.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the role of dual weighted residual methodology in this framework?",{"text":80,"@type":76},"It is used to derive computable upper bounds on the functional (approximation) error, enabling practical estimation rather than only qualitative guidance.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets and task are used to demonstrate the method?",{"text":84,"@type":76},"The method is evaluated on scientific datasets, including learning the observable-to-parameter map for the Navier-Stokes equation, where it improves generalization compared with existing architecture adaptation methods.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]