[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81828-en":3,"doc-seo-81828-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81828,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation","EO-Agents introduces a structured pipeline for Earth observation scientific hypothesis generation grounded in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical co-usage relations ranks candidate dataset pairings, while a three-agent LLM workflow filters, generates, and evaluates structured research hypotheses. On 1,475 NASA datasets, the system produces 160 hypotheses across domains such as ecohydrology, glaciology, aerosol–cloud interactions, vegetation phenology, and stratospheric chemistry. Predicted novel pairings achieve plausibility comparable to held-out real co-usages, indicating scientifically coherent yet unexplored combinations.","EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis  \nGeneration  \nMahyar Ghazanfari 1 Amin Tabrizian 1 Armin Mehrabian 2 3 Peng Wei 1  \narXiv :2607 .0 1584v 1 [ cs .AI] 2 Jul 2026  \nAbstract  \nLarge language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims. We present a pipeline for Earth observation that grounds hypothesis generation directly in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical cousage relations ranks candidate dataset pairings, and a three-agent LLM pipeline filters, generates, and evaluates structured research hypotheses. Applied to 1,475 NASA datasets, the system produces 160 hypotheses spanning multiple Earth-science domains, including ecohydrology, glaciology, aerosol–cloud interactions, vegetation phenology, and stratospheric chemistry. Modelpredicted novel dataset pairings are rated nearly as plausible as held-out real co-usages from the literature, indicating that the pipeline surfaces scientifically coherent yet unexplored combinations.  \nA 2 × 2 × 2 factorial experiment across GPT- 5.2 and Claude Sonnet 4.6 shows that hypothesis rankings remain stable, while absolute scores depend strongly on judge identity, highlighting limitations of single-judge LLM evaluation. Code, dataset, and generated hypotheses can be found here.  \n1. Introduction  \nEarth-observation (EO) research is fundamentally combinatorial. A typical study fuses a soil-moisture retrieval with a vegetation index, pairs a lidar canopy-height product with a global digital-elevation model, or cross-references atmospheric chemistry against stratospheric profiles. NASA  \nAccepted at ICML 2026 AI for Science Workshop. 1 George Washington University, Washington, DC, USA 2NASA Goddard Space Flight Center, Greenbelt, MD, USA 3ADNET Systems, Inc., Bethesda, MD, USA. Correspondence to: Mahyar Ghazanfari \u003C[mahyar.ghazanfari@gwu.edu](mahyar.ghazanfari@gwu.edu) >.  \nalone distributes over 8,000 EO datasets across twelve archives, spanning hundreds of instruments and five decades of observations. The scientific payoff is increasingly driven not by any single measurement but by the pairing a researcher chooses to study, and by whether that pairing has been attempted before. The combinatorial space of unordered pairings, ∼ 106 even under conservative filtering, far exceeds what any individual researcher can survey, and most of it remains scientifically unexplored.  \nRecent work has proposed large language models (LLMs) as engines for scientific hypothesis generation, producing research ideas in chemistry (Yang et al., 2025c ;b), biomedicine (Wang et al., 2024a), materials science (Ghafarollahi & Buehler, 2025), astrobiology (Saeedi et al., 2025), and machine-learning research itself (Si et al., 2025 ; Lu et al., 2024 ; Yamada et al., 2025) . The dominant paradigm grounds LLM generation in unstructured scientific text: a retriever surfaces related papers, a generator composes an idea from their content, and an evaluator scores the result (Baek et al., 2025 ; Yu et al., 2025) . This paradigm is effective for literature-driven domains where hypotheses are themselves textual claims; in observational, data-rich domains such as EO, however, a useful hypothesis is inseparable from the specific measurement products that would test it. Saying “use satellite data to study drought” is not a hypothesis; saying “combine SPL4SMGP soil moisture with MYD13Q1 vegetation indices to test whether 16-day EVI response depends on biome-specific soil-moisture thresholds”is. A literature-grounded pipeline can in principle reach the second statement by extracting dataset mentions from retrieved papers, but it must do so as a side-effect of freeform text generation. We argue that grounding the retrieval step directly in a typed knowledge graph of measurement products is a more direct route, because ever","cbCaijgCQXDSgNKo","https://ap.wps.com/l/cbCaijgCQXDSgNKo","pdf",1308423,6,1,23,"English","en",105,"# Abstract\n# Introduction\n## Motivation: combinatorial dataset pairing\n## Limitations of literature-grounded LLM paradigms\n## Structured retrieval from NASA EO knowledge graph\n## Three-agent LLM hypothesis pipeline\n## Experimental evaluation and findings","[{\"question\":\"How does EO-Agents ground hypothesis generation for Earth observation research?\",\"answer\":\"EO-Agents grounds hypotheses directly in the NASA Earth Observation Knowledge Graph by pinning each generated hypothesis to two named NASA datasets through structured retrieval and typed graph links.\"},{\"question\":\"What role does the heterogeneous GNN play in the pipeline?\",\"answer\":\"The heterogeneous graph neural network is trained on historical co-usage relations and ranks candidate dataset pairings by predicted co-usage likelihood, surfacing top pairs not observed in held-out co-usage splits.\"},{\"question\":\"How are hypotheses produced and judged in the three-agent LLM pipeline?\",\"answer\":\"A filter agent rescoring top candidates assesses plausibility and novelty, a generator agent formulates a structured hypothesis with question, testable claim, analysis method, and expected finding, and a judge agent rates importance, tractability, and novelty under blind and contextual settings.\"}]","EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation | PDF",1784176426,58,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"eo-agents-a-three-agent-llm-pipeline-for-earth-observation-hypothesis-generation","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/eo-agents-a-three-agent-llm-pipeline-for-earth-observation-hypothesis-generation/81828/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"How does EO-Agents ground hypothesis generation for Earth observation research?","Question",{"text":77,"@type":78},"EO-Agents grounds hypotheses directly in the NASA Earth Observation Knowledge Graph by pinning each generated hypothesis to two named NASA datasets through structured retrieval and typed graph links.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What role does the heterogeneous GNN play in the pipeline?",{"text":82,"@type":78},"The heterogeneous graph neural network is trained on historical co-usage relations and ranks candidate dataset pairings by predicted co-usage likelihood, surfacing top pairs not observed in held-out co-usage splits.",{"name":84,"@type":75,"acceptedAnswer":85},"How are hypotheses produced and judged in the three-agent LLM pipeline?",{"text":86,"@type":78},"A filter agent rescoring top candidates assesses plausibility and novelty, a generator agent formulates a structured hypothesis with question, testable claim, analysis method, and expected finding, and a judge agent rates importance, tractability, and novelty under blind and contextual settings.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]