[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86351-en":3,"doc-seo-86351-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86351,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Sense and Sensitivity Examining the Influence of Semantic Recall on Long Context Code Understanding","Large language models (LLMs) increasingly tackle long-code understanding, yet whether they capture operational semantics or depend on shortcuts remains unclear. The study distinguishes lexical recall (verbatim retrieval) from semantic recall (understanding what code does). Across 10 state-of-the-art LLMs, lexical recall stays near-perfect and position-independent, while semantic recall collapses when relevant code is centrally placed in long contexts. The paper introduces semantic recall sensitivity, a counterfactual line-removal method, and a new SemTrace task to expose severe positional failures.","Sense and Sensitivity: Examining the Influence of Semantic Recall on Long  \nContext Code Understanding  \nAdam Štorek Mukur Gupta Samira Hajizadeh Prashast Srivastava Suman Jana  \nColumbia University  \n{astorek, [suman}@cs.columbia.edu](suman}@cs.columbia.edu)  \n{mukur.gupta, sh4635, [ps3400}@columbia.edu](ps3400}@columbia.edu)  \narXiv :2505 . 13353v 5 [ cs .CL] 10 Jul 2026  \nAbstract  \nLarge language models (LLMs) are increasingly deployed for understanding large codebases, but whether they understand operational semantics of long code context or rely on pattern matching shortcuts remains unclear. We distinguish between lexical recall (retrieving code verbatim) and semantic recall (understanding operational semantics) . Evaluating  \n10 state-of-the-art LLMs, we find that while frontier models achieve near-perfect, positionindependent lexical recall, semantic recall degrades severely when code is centrally positioned in long contexts. We introduce semantic recall sensitivity to measure whether tasks require understanding of code’s operational semantics vs. permit pattern matching shortcuts. Through a novel counterfactual measurement method, we show that models rely heavily on pattern matching shortcuts to solve existing code understanding benchmarks. We propose a new task SemTrace, which achieves high semantic recall sensitivity through unpredictable operations; LLMs’ accuracy exhibits severe positional effects, with median accuracy drops of  \n92.73% versus CRUXEval’s 53 .36% as the relevant code snippet approaches the middle of the input code context. Our findings suggest current evaluations substantially underestimate semantic recall failures in long context code understanding.1  \n1 Introduction  \nLarge Language Models (LLMs) are increasingly applied to industry coding tasks (Edwards, 2024) that demand understanding of large codebases (Jimenez et al., 2024) . Recent advances (Dao et al., 2022 ; Peng et al., 2024 ; Su et al., 2024) enable these models to process extremely long inputs, up to millions of tokens (OpenAI, 2025) . However, a fundamental question remains unanswered:  \n1Our code is available at [https://github.com/](https://github.com/)[ ](https://github.com/)adamstorek/long-context-code-understanding.  \nwhen models solve code understanding tasks, are they processing the specific code provided in context, or applying memorized patterns from pretraining? This distinction becomes critical as LLMs are deployed in production environments where they must handle novel, project-specific code that cannot be solved through pattern matching alone.  \nPattern matching shortcuts may also result in LLMs missing subtle vulnerabilities (Ding et al., 2025) . We introduce a key distinction between two capabilities for code understanding in long contexts: lexical recall, meaning the ability to locate and reproduce code verbatim, and semantic recall, meaning the ability to remember what code does when it is run, i.e., its operational semantics (Winskel, 1993) . These capabilities are distinct; models can have perfect lexical recall yet fail at semantic recall, demonstrating they can access relevant code but not understand its effects. While needle-in-thehaystack (NIAH) benchmarks (Liu et al., 2024a,c) measure lexical recall, the relationship between lexical and semantic recall is not well understood.  \nMoreover, code understanding tasks like output prediction are designed to measure semantic recall, but can often be solved through pattern-matching shortcuts (recognizing familiar algorithms, applying memorized correlations) without requiring semantic recall of the specific implementation. This conflation poses a fundamental evaluation challenge: existing benchmarks may allow shortcuts that mask semantic recall failures. We introduce semantic recall sensitivity as a property of tasks: the degree to which solving the task requires semantically recalling specific code details rather than pattern matching.  \nTo investigate whether lexical an","cbCaio8fnanys8Gh","https://ap.wps.com/l/cbCaio8fnanys8Gh","pdf",541542,4,1,19,"English","en",105,"# Abstract\n# Introduction\n# Lexical vs Semantic Recall\n## Positional Effects in Long Contexts\n## Counterfactual Line-Removal Measurement\n# Semantic Recall Sensitivity\n# SemTrace Task","[{\"question\":\"What is the difference between lexical recall and semantic recall in long-context code understanding?\",\"answer\":\"Lexical recall is the ability to retrieve and reproduce code verbatim, while semantic recall is the ability to remember what the code does when executed (its operational semantics). The paper treats these as separable capabilities.\"},{\"question\":\"How do LLMs’ semantic recall and lexical recall change with code position in long contexts?\",\"answer\":\"Across 10 state-of-the-art LLMs, lexical recall remains near-perfect and largely position-independent. In contrast, semantic recall degrades severely when the relevant code is centrally positioned in long contexts.\"},{\"question\":\"Why does the paper claim existing benchmarks may underestimate semantic recall failures?\",\"answer\":\"Because benchmarks like CRUXEval can be solved using pattern-matching shortcuts. The proposed counterfactual method shows gradual degradation when code lines are removed, indicating low sensitivity to semantic recall requirements.\"}]",1784210738,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sense-and-sensitivity-examining-the-influence-of-semantic-recall-on-long-context-code-understanding","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/sense-and-sensitivity-examining-the-influence-of-semantic-recall-on-long-context-code-understanding/86351/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the difference between lexical recall and semantic recall in long-context code understanding?","Question",{"text":75,"@type":76},"Lexical recall is the ability to retrieve and reproduce code verbatim, while semantic recall is the ability to remember what the code does when executed (its operational semantics). The paper treats these as separable capabilities.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do LLMs’ semantic recall and lexical recall change with code position in long contexts?",{"text":80,"@type":76},"Across 10 state-of-the-art LLMs, lexical recall remains near-perfect and largely position-independent. In contrast, semantic recall degrades severely when the relevant code is centrally positioned in long contexts.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the paper claim existing benchmarks may underestimate semantic recall failures?",{"text":84,"@type":76},"Because benchmarks like CRUXEval can be solved using pattern-matching shortcuts. The proposed counterfactual method shows gradual degradation when code lines are removed, indicating low sensitivity to semantic recall requirements.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]