[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83168-en":3,"doc-seo-83168-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83168,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Hybrid Least Squares/Gradient Descent Methods for MIONets","Efficient hybrid least squares/gradient descent (LSGD) training is proposed for MIONets to accelerate optimization. The approach generalizes LSGD from DeepONets by treating MIONet as a multilinear function of the last-layer parameters across multiple branch networks and a trunk network. Alternating least squares solves least-squares systems for one branch at a time while factorizing large matrices using Kronecker products, Khatri-Rao products, and tensor permutation matrices. The method supports general L2 losses with regularization and optional linear operators acting on outputs in each loss term.","arXiv :2607 .06976v 1 [ cs .LG] 8 Jul 2026  \nHybrid Least Squares/Gradient Descent Methods for  \nMIONets  \nJun Choi 1 , Chang-Ock Lee 1 , and Minam Moon2  \n1 Department of Mathematical Sciences, KAIST, Daejeon 34141, KOREA  \n2 Department of Mathematics, Korea Military Academy, Seoul 01805, KOREA  \nAbstract  \nIn this paper, we propose an efficient hybrid least squares/gradient descent (LSGD) method for MIONets to accelerate training. This method generalizes the LSGD method for DeepONets. Since MIONet is the sum of the entrywise product of multiple branch networksand a trunk network, it can be viewed as a multilinear function with respect to the last layer parameters of each branch network. These sets of parameters can be optimized using the alternating least squares method, where we solve the LS system for a single branch network in turn. To handle the large-sized system matrix, we introduce Kronecker and Khatri-Rao products and tensor permutation matrices to factor the large matrix into small ones. Our method is compatible with a general type of L2 loss with regularization terms for the last layer parameters of each branch, where linear operators can be applied to the MIONet output in each loss term.  \nKey words. Hybrid least squares gradient descent method, MIONet, Kronecker product, Khatri-Rao product, tensor permutation matrix  \nMSC codes. 15A69, 47-08, 65F45, 65Y10, 68T07, 68T20  \n1 Introduction  \nThanks to the recent advances in scientific machine learning, the core architectures, including deep learning (DL) and deep neural networks (DNNs) have migrated to the field of scientific computing to enhance existing numerical methods for solving various partial differential equations (PDEs) . In particular, physics-informed neural network (PINN) [25] is the most successful and widely used method, where PINN represents the solution of PDE as a DNN and finds the solution by training the DNN using a physics-informed loss (PI-loss) with the automatic differentiation method [1] . However, since PINN requires separate training for different PDE instances, the need for a mapping between components and solutions of PDEs using DL architecture has emerged, which has now been generalized to neural operator mapping between function spaces. There are various examples of neural operators, including Deep Operator Network (DeepONet) [19], Fourier Neural Operator [17], Graph Kernel Network [18], PCA-based Model Reduction [2], and Multi-Wavelet Neural Operator [8] .  \nAmong these neural operators, DeepONet is the most widely used framework for neural operators, which possesses the universal approximation property. It consists of the inner product  \nFunding: This work was supported by Basic Science Research Program through the National Research Foundation (NRF) of Korea funded by the Ministry of Education [RS2025–25397599] .  \nof outputs from two neural networks, branch and trunk, where the branch network encodes input functions and the trunk network encodes coordinates of the output function domain. Based on the DeepONet architecture, many variants have been proposed, such as POD-DeepONet [20], Multifidelity DeepONet [21], NOMAD [29], Multiple-Input Operator Network (MIONet) [13], Shift-DeepONet [9], HyperDeepONet [16], and Geom-DeepONet [10] .  \nIn this paper, we focus on MIONet since it is a direct generalization of DeepONet, which maps several input functions to a single output function with the corresponding universal approximation theorem (UAT) [13, Theorem 3.1] . Instead of a single branch network in DeepONet, MIONet uses multiple branch networks to encode each input function and computes the entrywise product of the outputs from each branch to perform an inner product with the output of the trunk network.  \nHowever, the training for MIONet is challenging because the entrywise product and inner product among several networks make the structure more complex, and a sufficiently large dataset is needed for meaningful training. This makes the ","cbCaiorub72WCcTW","https://ap.wps.com/l/cbCaiorub72WCcTW","pdf",909824,1,20,"English","en",105,"# Introduction\n## Neural operators and DeepONet framework\n## Motivation for MIONet training challenges\n## Proposed hybrid LSGD/ALS method and matrix factorization","[{\"question\":\"What does the proposed hybrid LSGD method optimize for MIONets?\",\"answer\":\"It optimizes the last-layer parameters of each branch network in MIONets by formulating a minimization of sums of squared multilinear functions under a general L2 loss with optional regularization.\"},{\"question\":\"How does the method handle the alternating optimization across branch networks?\",\"answer\":\"It fixes the last-layer parameters of all but one branch, then solves a least-squares system for the unfixed branch. By alternating which branch is unfixed, all branch last-layer parameters are optimized in sequence.\"},{\"question\":\"Why are Kronecker and Khatri-Rao products used in the approach?\",\"answer\":\"They factor the large least-squares system matrix into smaller component matrices derived from branch and trunk networks, making the otherwise large linear system computationally tractable.\"}]",1784185725,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"hybrid-least-squaresgradient-descent-methods-for-mionets","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hybrid-least-squaresgradient-descent-methods-for-mionets/83168/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the proposed hybrid LSGD method optimize for MIONets?","Question",{"text":75,"@type":76},"It optimizes the last-layer parameters of each branch network in MIONets by formulating a minimization of sums of squared multilinear functions under a general L2 loss with optional regularization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method handle the alternating optimization across branch networks?",{"text":80,"@type":76},"It fixes the last-layer parameters of all but one branch, then solves a least-squares system for the unfixed branch. By alternating which branch is unfixed, all branch last-layer parameters are optimized in sequence.",{"name":82,"@type":73,"acceptedAnswer":83},"Why are Kronecker and Khatri-Rao products used in the approach?",{"text":84,"@type":76},"They factor the large least-squares system matrix into smaller component matrices derived from branch and trunk networks, making the otherwise large linear system computationally tractable.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":28,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]