[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83408-en":3,"doc-seo-83408-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83408,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation","Repository-level code generation must implement target functions while respecting cross-file dependencies and project conventions, yet common retrieval methods largely optimize lexical, structural, or semantic similarity. ProjAgent introduces procedural similarity as an explicit retrieval signal by decomposing the target function into intermediate reasoning steps and using an agentic workflow to retrieve repository functions with matching procedural behavior at each step. The retrieved procedural context is fused with semantic retrieval and refined via conservative static-analysis feedback, improving Pass@1 to 41.14% on REPOCOD.","arXiv :2607 .0869 1v 1 [ cs . SE] 9 Jul 2026  \nProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation  \nQIHONG CHEN, University of California, Irvine, USA AARON IMANI✉ , University of California, Irvine, USA IFTEKHAR AHMED, University of California, Irvine, USA  \nRepository-level code generation requires implementing target functions while accounting for complex crossfile dependencies and project-specific conventions. Existing retrieval methods predominantly rely on lexical, structural, or semantic similarity, often overlooking repository functions that implement similar procedural logic despite differing in identifiers or application domains. We propose ProjAgent, a repository-level code generation system that introduces procedural similarity as an explicit retrieval signal. ProjAgent decomposes the target function into intermediate reasoning steps and employs an agentic workflow to retrieve repository functions that exhibit similar procedural behavior at each step. The retrieved procedural context is integrated with conventional semantic retrieval to construct a richer repository context for code generation. ProjAgent further incorporates a conservative static-analysis feedback loop that iteratively repairs generated code using compiler and static-analysis feedback. Evaluated on REPOCOD, ProjAgent achieves 41.14% Pass@1, outperforming existing retrieval-based baselines. These results demonstrate that procedural similarity is an effective and previously unexplored retrieval dimension for repository-level code generation.  \n1 Introduction  \nLarge Language Models (LLMs) have demonstrated strong performance across a wide range of Software Engineering (SE) tasks. From automated bug fixing to test case generation, LLM-based approaches have achieved promising results on established benchmarks such as SWE-bench [17] and BigCodeBench [49]. These advances have accelerated the adoption of LLMs in software development workflows, supporting activities including code writing, code review, and software maintenance. Among these applications, code generation has attracted particular attention due to its broad practical utility and increasing adoption in real-world development [37, 50] .  \nDespite these advances, repository-level code generation remains substantially more challenging than standalone code generation [3, 6, 41] . Unlike standalone benchmarks, real-world software development rarely involves implementing isolated functions. Instead, developers work within repositories where functions depend on utilities, type definitions, APIs, and project-specific conventions distributed across multiple files [24, 44] . When relevant repository context is unavailable, LLMs frequently hallucinate APIs, invoke nonexistent functions, or generate implementations that violate project conventions, resulting in incorrect or non-executable code [21] . Prior work has shown that removing repository context consistently degrades code generation performance across models [24]. Consequently, repository-level code generation depends not only on the generation capability of an LLM, but also on its ability to retrieve repository context that is useful for solving the target programming task.  \nExisting context retrieval methods for repository-level code generation, such as BM25 and dense embedding search, rely primarily on lexical or semantic similarity [12, 36, 44] . These methods were originally developed for code search, where retrieving surface-level similar examples is often sufficient [21]. In repository-level code generation, however, critical context for the target function (i.e., the function to be generated) may come from helper functions that it depends on, even when those functions differ substantially in naming, data types, or domain vocabulary. Useful context  \nAuthors’ Contact Information: Qihong Chen, [chenqh@uci.edu](chenqh@uci.edu), University of California, Irvine, Irvine, USA; Aaron Imani, University of California, I","cbCaidvGDi3BNE2e","https://ap.wps.com/l/cbCaidvGDi3BNE2e","pdf",1392616,4,1,19,"English","en",105,"# Introduction\n## Motivation for procedural similarity in repository-level generation\n## Limitations of existing retrieval signals\n## ProjAgent overview and methodology","[{\"question\":\"Why is repository-level code generation harder than standalone function generation?\",\"answer\":\"Repository-level generation must integrate cross-file dependencies, helper utilities, type definitions, APIs, and project-specific conventions. Without useful repository context, LLMs often hallucinate APIs or produce non-conforming, non-executable code.\"},{\"question\":\"What gap do existing retrieval methods have for this task?\",\"answer\":\"Existing context retrieval methods like BM25 and dense embeddings primarily use lexical or semantic similarity. They can miss helper functions that implement similar reasoning steps but differ in identifiers, data types, or domain vocabulary, leading to degraded generation quality.\"},{\"question\":\"How does ProjAgent use procedural similarity to improve code generation?\",\"answer\":\"ProjAgent decomposes the target function into intermediate reasoning steps, retrieves repository functions that exhibit similar procedural behavior at each step through an agentic workflow, and integrates this procedural context with conventional semantic retrieval. It then iteratively repairs generated code using compiler and static-analysis feedback.\"}]",1784187358,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"projagent-procedural-similarity-retrieval-for-repository-level-code-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/projagent-procedural-similarity-retrieval-for-repository-level-code-generation/83408/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is repository-level code generation harder than standalone function generation?","Question",{"text":75,"@type":76},"Repository-level generation must integrate cross-file dependencies, helper utilities, type definitions, APIs, and project-specific conventions. Without useful repository context, LLMs often hallucinate APIs or produce non-conforming, non-executable code.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What gap do existing retrieval methods have for this task?",{"text":80,"@type":76},"Existing context retrieval methods like BM25 and dense embeddings primarily use lexical or semantic similarity. They can miss helper functions that implement similar reasoning steps but differ in identifiers, data types, or domain vocabulary, leading to degraded generation quality.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ProjAgent use procedural similarity to improve code generation?",{"text":84,"@type":76},"ProjAgent decomposes the target function into intermediate reasoning steps, retrieves repository functions that exhibit similar procedural behavior at each step through an agentic workflow, and integrates this procedural context with conventional semantic retrieval. It then iteratively repairs generated code using compiler and static-analysis feedback.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]