[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83656-en":3,"doc-seo-83656-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83656,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Bringing Agentic Search to Earth Observation Data Discovery","NASA and its data centers provide thousands of geoscience datasets and tools, yet locating the best resources for a specific research question remains difficult even for domain experts. The work introduces an agentic search system available as a public service, which converts natural-language research queries into matched datasets and tools. It leverages NASA EO-KG, builds NASA-EOBench with 47k query–dataset pairs, and trains a neural scorer that improves retrieval beyond cosine and BM25. Score fusion further boosts Recall@10 and MRR, and a zero-shot agentic reranking stage adds additional MRR gains by combining LLM reasoning with supervised retrieval.","Bringing Agentic Search to Earth Observation Data  \nDiscovery  \nMinghan Yu∗  \nUniversity of Maryland, College Park [my6489@umd.edu](my6489@umd.edu)  \nYouran Sun∗  \nUniversity of Maryland, College Park [sun1245@umd.edu](sun1245@umd.edu)  \narXiv :2607 .02387v 1 [ cs .IR] 2 Jul 2026  \nChugang Yi  \nUniversity of Maryland, College Park [chugang@umd.edu](chugang@umd.edu)  \nYixin Wen† University of Florida [yixin.wen@ufl.edu](yixin.wen@ufl.edu)  \nHaizhao Yang†  \nUniversity of Maryland, College Park  \n[hzyang@umd.edu](hzyang@umd.edu)  \n∗Equal contribution. †Corresponding authors.  \nAbstract  \nNASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search system, deployed as a public service for the geoscience community, that takes a natural-language research query and returns the matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic search. From the NASA Earth Observation Knowledge Graph (NASA EO-KG) we derive NASA-EOBench, an open benchmark of 47k query–dataset pairs (21k task-based queries) . A neural scorer fine-tuned on NASA-EO-Bench beats cosine and BM25 baselines.  \nFurther combining it with BM25 via score fusion raises both Recall@10 (R@10) and MRR by over 5 × . On top of this supervised pipeline, we add a zero-shot agentic reranking stage that, without any additional training, lifts MRR by 28% on a stratified N=200 subset, showing that LLM reasoning is complementary to supervised retrieval.  \n1 Introduction  \nEarth-observation data discovery is less a problem of data scarcity than of navigating fragmented metadata, tools, and access pathways. NASA and its affiliated data centers host thousands of datasets across dozens of Distributed Active Archive Centers (DAACs), together with tools such as Worldview, Giovanni, Science Discovery Engine, and Harmony.1 This fragmentation makes it hard even for domain experts to locate the data that best matches their own research question.  \n1Worldview [https://worldview.earthdata.nasa.gov/](https://worldview.earthdata.nasa.gov/), Giovanni [https://giovanni](https://giovanni) . [gsfc.nasa.gov/giovanni/](gsfc.nasa.gov/giovanni/), Science Discovery Engine [https://science.data.nasa.gov/](https://science.data.nasa.gov/)[ ](https://science.data.nasa.gov/)science-discovery-engine/, Harmony [https://harmony.earthdata.nasa.gov/](https://harmony.earthdata.nasa.gov/) .  \nPreprint. Under review.  \nFigure 1: Overview of the three-stage agentic search pipeline (Section 5.1) . Stage 1: the Router first attempts to resolve each query via NASA official tools (Harmony, SDE, WorldView, Giovanni); queries that can be fully answered here terminate early. Stage 2: if no official tool suffices, the hybrid BM25 + NN-SSC retriever surfaces the most relevant datasets from the NASA CMR corpus. Stage 3 : retrieved candidates are reranked by an LLM that autonomously invokes web search and arXiv paper lookup to ground its ranking decisions in external context; the final ranked list is returned to the user.  \nA large language model (LLM) offers researchers a natural interface for expressing dataset-retrieval intent in natural language. In a highly domain-specific setting such as geoscience dataset retrieval, however, using an LLM directly for queries faces two cumulative challenges. First, the pre-training corpora of general LLMs are dominated by general web text and lack the observational data and domain knowledge of geoscience, so domain queries are often neither accurately understood nor directly answerable in a trustworthy way. Second, even when Retrieval-augmented generation (RAG) Lewis et al. [2020] injects retrieved evidence at query time to close this gap, the LLM’s context window still imposes a hard limit on the number of candidat","cbCaipNHRoK0PsdA","https://ap.wps.com/l/cbCaipNHRoK0PsdA","pdf",522450,3,1,19,"English","en",105,"# Introduction\n## Challenges in Geoscience Data Discovery\n## Agentic Search Pipeline Overview\n## Trustable and Verifiable Evaluation Criteria\n# Method and Benchmark Construction","[{\"question\":\"Why is earth-observation data discovery difficult even for domain experts?\",\"answer\":\"Because NASA and related centers fragment datasets across many DAACs along with multiple metadata and access tools, making it hard to navigate and match resources to a specific research question.\"},{\"question\":\"What is the proposed agentic search system designed to do?\",\"answer\":\"It takes a natural-language research query and returns the most matching datasets and tools, using an agentic, multi-stage pipeline that includes routing, hybrid retrieval, and LLM-based reranking with external evidence grounding.\"},{\"question\":\"How is NASA-EO-Bench constructed and how does it support evaluation?\",\"answer\":\"NASA-EOBench is derived from edges of the NASA Earth Observation Knowledge Graph and released as an open benchmark containing 47k query–dataset pairs, enabling citation-grounded, protocol-consistent evaluation using citation-based metrics.\"}]",1784189555,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"bringing-agentic-search-to-earth-observation-data-discovery","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/bringing-agentic-search-to-earth-observation-data-discovery/83656/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is earth-observation data discovery difficult even for domain experts?","Question",{"text":75,"@type":76},"Because NASA and related centers fragment datasets across many DAACs along with multiple metadata and access tools, making it hard to navigate and match resources to a specific research question.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed agentic search system designed to do?",{"text":80,"@type":76},"It takes a natural-language research query and returns the most matching datasets and tools, using an agentic, multi-stage pipeline that includes routing, hybrid retrieval, and LLM-based reranking with external evidence grounding.",{"name":82,"@type":73,"acceptedAnswer":83},"How is NASA-EO-Bench constructed and how does it support evaluation?",{"text":84,"@type":76},"NASA-EOBench is derived from edges of the NASA Earth Observation Knowledge Graph and released as an open benchmark containing 47k query–dataset pairs, enabling citation-grounded, protocol-consistent evaluation using citation-based metrics.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]