[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85226-en":3,"doc-seo-85226-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85226,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","LLM-PDESR Robust PDE Discovery via Subdomain Weighted Residuals and LLM-Guided Symbolic Hypothesis Generation","Discovering governing partial differential equations (PDEs) from noisy observations is a core scientific machine learning problem. Conventional symbolic regression methods often fail in huge combinatorial search spaces and cannot effectively inject domain priors, while pointwise residuals and discrete finite differences amplify high-frequency noise and distort optimization feedback. LLM-PDESR combines LLM-guided structural hypothesis generation with a rigorous evaluation pipeline using C4 quintic splines for robust differentiation and subdomain weighted residuals as low-pass filtering. A Pareto-driven loop refines candidates to balance accuracy and parsimony. Experiments on 23 canonical PDEs plus five novel systems, including recovery from noisy ERA5 reanalysis, show strong gains in structural recovery, noise resilience, and reduced spurious complexity.","arXiv :2607 . 10546v 1 [ cs .LG] 12 Jul 2026  \nLLM-PDESR: Robust PDE Discovery via Subdomain Weighted Residuals and LLM-Guided Symbolic Hypothesis Generation  \nJinyang Du 1 Hao Ma 1 Xiaohu Shi2 Bo Yang 1 ,3 Yanchun Liang4 Heow Pueh Lee5 Chunguo Wu 1 ,3 ,∗  \n1 College of Computer Science and Technology, Jilin University, Changchun, China  \n2 School of Big Data and Artificial Intelligence, Guangdong University of Finance and Economics  \nGuangzhou, China  \n3Key Laboratory of Symbolic Computation and Knowledge Engineering  \nof Ministry of Education, Jilin University, Changchun, China  \n4 School of Computer Science, Zhuhai College of Science and Technology, Zhuhai, China  \n5Department of Mechanical Engineering, National University of Singapore, Singapore  \n∗ Corresponding author: [wucg@jlu.edu.cn](wucg@jlu.edu.cn)  \nAbstract  \nDiscovering governing partial differential equations (PDEs) from noisy observational data is a fundamental challenge in scientific machine learning. Traditional symbolic regression (SR) methods often struggle to identify accurate equations within vast combinatorial search spaces, largely due to their inability to incorporate essential domain-specific prior knowledge. Furthermore, reliance on pointwise evaluations and discrete finite differences inherently amplifies high-frequency noise, creating deceptive fitness landscapes that derail the optimization process. To resolve these bottlenecks, we propose LLM-PDESR, a framework that integrates the structural hypothesis generation of Large Language Models (LLMs) with a mathematically rigorous evaluation environment. By employing C4 continuous quintic splines for robust differentiation and subdomain weighted residuals as natural lowpass filters, our approach effectively mitigates the fitness landscape distortion that plagues existing methods. A Pareto-driven feedback loop then enables the LLM to iteratively refine candidate equations, balancing predictive accuracy with structural parsimony. We evaluate LLM-PDESR on 23 canonical PDEs and five structurally novel equations (including a multivariate system) specifically designed to preclude dataset memorization and test true discovery capabilities. Demonstrating real-world applicability, the framework successfully extracts a consistent structural skeleton for an interpretable 1D dynamical surrogate (1D-CACE) directly from noisy ERA5 reanalysis data. Extensive experiments and out-of-distribution testing confirm that LLM-PDESR significantly outperforms state-of-the-art methodologies in structural recovery, noise resilience, and the avoidance of spurious complexity and equation bloat.  \n1 Introduction  \nA central pursuit in modern computational science is bridging the gap between abundant empirical measurements and the underlying mathematical laws that dictate system dynamics [1–5] . While the formalization of this problem through symbolic regression (SR) has spurred significant advances [6–8], traditional data-driven methodologies ranging from sparse regression [9, 10] to genetic programming [11–13] exhibit critical limitations. Constrained by predefined function libraries, these approaches  \nPreprint.  \nstruggle to scale across vast combinatorial hypothesis spaces and lack the architectural flexibility to incorporate domain-specific physical priors [14] .  \nRecently, Large Language Models (LLMs) have emerged as highly promising engines for this task [15–19] . By treating equations as executable structures, LLMs leverage their pre-trained scientific knowledge to propose physically plausible structural hypotheses [20] . However, extending these LLM-driven paradigms to the discovery of complex, time-varying PDEs is bottlenecked by severe numerical instabilities. To compute the necessary spatial derivatives for evaluating PDE hypotheses, traditional optimization frameworks predominantly rely on standard finite difference (FD) methods and pointwise residual evaluations [21–23] . These discrete operations inherently amplify h","cbCaipXKI5sAZzg7","https://ap.wps.com/l/cbCaipXKI5sAZzg7","pdf",14543468,2,1,28,"English","en",105,"# Introduction\n## Motivation and limitations of existing symbolic regression\n## LLM-guided equation hypothesis generation and numerical bottlenecks\n## The proposed LLM-PDESR framework\n### C4 continuous quintic spline smoothing\n### Subdomain weighted residual evaluation\n### Pareto-driven feedback and model parsimony\n# Evaluation and results\n## Benchmarks with canonical and structurally novel PDEs\n## Real-world case: extracting a 1D dynamical surrogate from ERA5","[{\"question\":\"What makes PDE discovery from noisy data especially challenging for traditional symbolic regression?\",\"answer\":\"Traditional symbolic regression struggles with vast combinatorial search spaces and limited ability to use domain-specific physical priors. Pointwise residual evaluation with discrete finite differences further amplifies high-frequency noise, producing misleading fitness landscapes that derail optimization.\"},{\"question\":\"How does LLM-PDESR improve numerical stability during PDE hypothesis evaluation?\",\"answer\":\"LLM-PDESR replaces noise-sensitive finite-difference differentiation with C4-continuous quintic spline smoothing, yielding noise-robust derivatives up to fourth order. It also replaces pointwise residuals with subdomain weighted residuals, which integrate residuals against smooth test functions to act as a natural low-pass filter.\"},{\"question\":\"How is the LLM guided to balance equation accuracy and complexity in LLM-PDESR?\",\"answer\":\"A Pareto-driven feedback loop iteratively refines candidate equations by balancing predictive accuracy with structural parsimony. This mechanism helps prune unphysical complexity and equation bloat while improving structural recovery.\"}]",1784201854,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"llm-pdesr-robust-pde-discovery-via-subdomain-weighted-residuals-and-llm-guided-symbolic-hypothesis-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/llm-pdesr-robust-pde-discovery-via-subdomain-weighted-residuals-and-llm-guided-symbolic-hypothesis-generation/85226/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes PDE discovery from noisy data especially challenging for traditional symbolic regression?","Question",{"text":75,"@type":76},"Traditional symbolic regression struggles with vast combinatorial search spaces and limited ability to use domain-specific physical priors. Pointwise residual evaluation with discrete finite differences further amplifies high-frequency noise, producing misleading fitness landscapes that derail optimization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LLM-PDESR improve numerical stability during PDE hypothesis evaluation?",{"text":80,"@type":76},"LLM-PDESR replaces noise-sensitive finite-difference differentiation with C4-continuous quintic spline smoothing, yielding noise-robust derivatives up to fourth order. It also replaces pointwise residuals with subdomain weighted residuals, which integrate residuals against smooth test functions to act as a natural low-pass filter.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the LLM guided to balance equation accuracy and complexity in LLM-PDESR?",{"text":84,"@type":76},"A Pareto-driven feedback loop iteratively refines candidate equations by balancing predictive accuracy with structural parsimony. This mechanism helps prune unphysical complexity and equation bloat while improving structural recovery.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]