[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82225-en":3,"doc-seo-82225-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82225,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Evaluating Semantic and Quality-Aware Retrieval for Source Code Repositories","Keyword-based retrieval in source code repositories underperforms when queries are written in natural language or target implementation intent and code quality instead of exact tokens. The work evaluates a prototype retrieval system combining function-level fragmentation, text-and-code embeddings stored in ChromaDB, LLM-derived quality metadata, and four retrieval modes: semantic, quality-filtered, hybrid, and automatic routing. Experiments use an educational C-code corpus with 15 manually judged queries, reporting semantic retrieval as the strongest overall approach.","Research Article Open Access  \nMarek Horváth* and Emília Pietriková  \nEvaluating Semantic and Quality-Aware Retrieval for Source Code Repositories  \narXiv :2607 .09 16 1v 1 [ cs . SE] 10 Jul 2026  \nAbstract: Keyword-based retrieval is limited for sourcecode repositories when queries are expressed in natural language or concern implementation intent and code quality rather than exact tokens. This study evaluates a prototype retrieval system that combines function-level fragmentation, text-and-code embeddings, ChromaDB vector storage, LLM-derived quality metadata, and four retrieval modes: semantic, quality-filtered, hybrid, and automatic routing. The concrete evaluation uses an educational Ccode corpus. The full corpus contains 563 anonymized programmer identifiers and 8,951 C files; a reproducible 10% indexed sample contains 56 programmer identifiers, 847 files, and 3,839 fragments. Across 15 manually judged queries, semantic retrieval achieved nDCG@5 of 0.820, Success@5 of 0.800, and MRR of 0.644 . The automatic router selected the expected mode for all 15 queries. In a small manual audit, LLM-derived quality scores were within one point of the manual assessment for 9 of 12 fragments. Within the reported query set, semantic retrieval was the strongest overall mode, while explicit quality metadata was most useful for explicitly quality-oriented queries.  \nKeywords: semantic code search, retrieval-augmented generation, vector embeddings, source code analysis, code quality evaluation  \n1 Introduction  \nSource-code repositories are difficult to inspect when the search intent concerns behavior, design choices, or code quality rather than exact tokens. Exact-token search remains useful when the user knows a function name, library call, or identifier, but it is less suitable for intent-level questions such as which fragments read from files, use dynamic memory allocation, or handle input errors. Such questions  \n*Corresponding author: Marek Horváth, Technical University of Košice, Faculty of Electrical Engineering and Informatics, Department of Computers and Informatics, Košice, Slovakia, E-mail: [marek.horvath@tuke.sk](marek.horvath@tuke.sk)  \nEmília Pietriková, Technical University of Košice, Faculty of Electrical Engineering and Informatics, Department of Computers and Informatics, Košice, Slovakia, E-mail:  \n[emilia.pietrikova@tuke.sk](emilia.pietrikova@tuke.sk)  \nrequire retrieval based on behavior and context rather than literal token overlap. Related source-code exploration work has also used static analysis and programmer profiles to summarize code artifacts beyond direct keyword matching [15] .  \nNatural-language code search addresses part of this problem by mapping queries and code fragments into a shared vector space [13] . In source-code repositories, however, topical similarity is not the only useful signal. Users may also ask quality-oriented questions about readability, error handling, input validation, or memory management. Adding explicit quality metadata can support such questions, but it can also distort topical relevance by promoting fragments that score well on quality dimensions while only weakly matching the requested functionality.  \nEducational C-code repositories provide a useful concrete evaluation setting because many implementations solve related tasks under comparable constraints. Such repositories occur in programming courses that combine repeated assignments, automated assessment, and gamecreative or problem-based learning activities [16, 17] . Atthe same time, findings from such a corpus should not be generalized to other repository types without further evaluation. This study therefore compares retrieval modes rather than assuming that quality scoring necessarily improves search. The evaluated prototype indexes sourcecode fragments, attaches LLM-derived quality metadata, stores embeddings and metadata in ChromaDB, and supports semantic, quality-filtered, hybrid, and automatically routed retrieval. R","cbCaiiaevdYBUVsx","https://ap.wps.com/l/cbCaiiaevdYBUVsx","pdf",426903,2,1,12,"English","en",105,"# Introduction\n# Background","[{\"question\":\"Why keyword-based retrieval is limited for source code repositories?\",\"answer\":\"It struggles when queries describe natural-language intent or code-quality aspects rather than exact tokens. This makes literal token overlap a weak signal for behavior and implementation intent.\"},{\"question\":\"What retrieval system components and modes does the study evaluate?\",\"answer\":\"The prototype fragments functions, uses text-and-code embeddings in ChromaDB, attaches LLM-derived quality metadata, and supports semantic, quality-filtered, hybrid, and automatic routing retrieval modes.\"},{\"question\":\"What do the reported results suggest about semantic vs. quality-aware retrieval?\",\"answer\":\"Across the reported query set, semantic retrieval shows the strongest overall performance. Explicit quality metadata is most useful for queries that are explicitly oriented toward quality dimensions.\"}]",1784178962,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-semantic-and-quality-aware-retrieval-for-source-code-repositories","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evaluating-semantic-and-quality-aware-retrieval-for-source-code-repositories/82225/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why keyword-based retrieval is limited for source code repositories?","Question",{"text":75,"@type":76},"It struggles when queries describe natural-language intent or code-quality aspects rather than exact tokens. This makes literal token overlap a weak signal for behavior and implementation intent.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What retrieval system components and modes does the study evaluate?",{"text":80,"@type":76},"The prototype fragments functions, uses text-and-code embeddings in ChromaDB, attaches LLM-derived quality metadata, and supports semantic, quality-filtered, hybrid, and automatic routing retrieval modes.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the reported results suggest about semantic vs. quality-aware retrieval?",{"text":84,"@type":76},"Across the reported query set, semantic retrieval shows the strongest overall performance. Explicit quality metadata is most useful for queries that are explicitly oriented toward quality dimensions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]