[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84831-en":3,"doc-seo-84831-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84831,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark","Scientific research increasingly depends on open-source software, yet finding relevant scientific tools among over 600 million GitHub repositories remains difficult. Existing code search benchmarks emphasize general software engineering and miss scientific computing’s domain-specific terminology and data formats. This work releases a curated corpus of 5,264 high-quality, domain-classified NASA scientific repositories enriched with cleaned READMEs and extracted context. It adds two retrieval benchmarks: 219 expert-built repository queries and a large code snippet benchmark with 117,950 snippets and 119,720 queries across seven languages, with public HuggingFace release.","arXiv :2607 .05443v 1 [ cs .IR] 3 Jul 2026  \nScientific Code Search at Scale: A Multi-Domain Dataset and Benchmark  \nNishan Pantha 1 , Pranath Reddy Kumbam 1 , Sajil Awale 1 , Pushwitha Krishnappa 1 , Muthukumaran Ramasubramanian 1 , Nidhi Jha 1 , Emily Foshee 1 , Ankur Kumar 1 , Rachel Slank2 , Ashkbiz Danehkar2 , and Rahul Ramachandran3  \n1 The University of Alabama in Huntsville (UAH), Huntsville, AL, USA  \n2 Universities Space Research Association (USRA), USA  \n3 NASA Marshall Space Flight Center, Huntsville, AL, USA  \nAbstract  \nScientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 million GitHub repositories remains challenging. Existing code search benchmarks focus on general software engineering tasks and fail to capture the domain-specific vocabulary and needs of scientific computing. We present a curated corpus of 5,264 high-quality, domain-classified scientific repositories spanning five NASA Science Mission Directorate divisions—Earth Science, Astrophysics, Planetary Science, Heliophysics, and Biological & Physical Sciences—enriched with cleaned READMEs, extracted topics, and additional context from crawled links. Building on this corpus, we introduce two novel information retrieval benchmarks: (1) a repository search benchmark with 219 expert-curated queries designed by domain scientists, and (2) a large-scale code snippet retrieval benchmark containing 117,950 code snippets and 119,720 queries across seven programming languages. Baseline evaluations on repository search reveal significant performance variation across scientific domains. Code snippet retrieval proves equally challenging, with substantial variation driven by differing documentation practices, coding standards, and programming language conventions across scientific communities. All datasets and benchmarks are publicly released on HuggingFace to support research on scientific tool discovery.  \nKeywords: code search, information retrieval, scientific software, benchmark datasets, code retrieval  \n1 Introduction  \nThe open science movement has transformed how scientific software is developed and shared [18] . Researchers increasingly publish code on GitHub, enabling reproducibility and accelerating discovery through reuse [3 , 6] . However, with over 600 million repositories on GitHub 1 , finding relevant scientific tools remains challenging. Scientific code discovery differs from general-purpose code search: researchers seek tools for specific data formats (e.g. , FITS in astronomy, NetCDF in climate science), particular instruments or missions, and domain-specific algorithms—queries requiring specialized vocabulary that general search systems struggle to capture [6] .  \n1[https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-secon](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-secon)d-as-ai-leads-typescript-to-1/  \nCurrent platforms like GitHub rely on lexical matching rather than semantic understanding [11]2. A researcher searching for “software to analyze time-series observations from the Kepler mission”may find nothing, while the same tool indexed under “exoplanet transit photometry pipeline” remains undiscovered. This discoverability gap leads to duplicated effort [10], hinders reproducibility [7], and confines effective tools to their originating communities. Existing code search benchmarks—CodeSearchNet [11], CoSQA [9], AdvTest [15]—draw from general-purpose software, reflect software engineering rather than research queries, and lack scientific vocabulary. Progress requires evaluation infrastructure grounded in real scientific queries and expert relevance judgments.  \nWe present datasets and benchmarks for scientific code discovery. Our contributions include:  \n1. Curated Scientific Repository Corpus: 5,264 high-quality, domain-classified GitHub repositories spanning five NASA SMD divis","cbCain1BS2YE1UMo","https://ap.wps.com/l/cbCain1BS2YE1UMo","pdf",622353,3,1,32,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Data Collection","[{\"question\":\"Why is scientific code search harder than general code search?\",\"answer\":\"Scientific discovery requires domain-specific vocabulary, data formats, instruments, and mission-related algorithms. General lexical search benchmarks and systems struggle to match these specialized queries.\"},{\"question\":\"What dataset and benchmarks are introduced in this work?\",\"answer\":\"It provides a curated corpus of 5,264 domain-classified scientific repositories and two retrieval benchmarks: a repository search benchmark with 219 expert-curated queries and a code snippet retrieval benchmark with 117,950 snippets and 119,720 queries across seven programming languages.\"},{\"question\":\"How do the authors evaluate retrieval across scientific domains?\",\"answer\":\"Baseline evaluations for repository search show significant performance variation across scientific domains, with code snippet retrieval also varying due to documentation practices, coding standards, and language conventions.\"}]",1784198593,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scientific-code-search-at-scale-a-multi-domain-dataset-and-benchmark","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/scientific-code-search-at-scale-a-multi-domain-dataset-and-benchmark/84831/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is scientific code search harder than general code search?","Question",{"text":75,"@type":76},"Scientific discovery requires domain-specific vocabulary, data formats, instruments, and mission-related algorithms. General lexical search benchmarks and systems struggle to match these specialized queries.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What dataset and benchmarks are introduced in this work?",{"text":80,"@type":76},"It provides a curated corpus of 5,264 domain-classified scientific repositories and two retrieval benchmarks: a repository search benchmark with 219 expert-curated queries and a code snippet retrieval benchmark with 117,950 snippets and 119,720 queries across seven programming languages.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the authors evaluate retrieval across scientific domains?",{"text":84,"@type":76},"Baseline evaluations for repository search show significant performance variation across scientific domains, with code snippet retrieval also varying due to documentation practices, coding standards, and language conventions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]