[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82016-en":3,"doc-seo-82016-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82016,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Jet-Long Efficient Long-Context Extension with Dynamic Bifocal RoPE","Modern LLMs need long-context support for retrieval-augmented generation, repository-level coding, and agentic workflows where accumulated reasoning and tool traces exceed the pretraining window. Zero-shot context extension dominates for open-weight checkpoints, but existing methods rely on a single fixed or fitted rescaling schedule that trades short-context fidelity for long-context stability. Jet-Long introduces a tuning-free bifocal RoPE with an analytic, parameter-free dynamic schedule, preserving exact behavior inside the native window while extrapolating cleanly to long sequences with minimal inference overhead.","arXiv :2607 .07740v2 [ cs .LG] 10 Jul 2026  \nJet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE  \nHaozhan Tang, Zerui Wang, Yuxian Gu, Song Han, Han Cai  \nNVIDIA  \n[https://github.com/jet-ai-projects/jet-long](https://github.com/jet-ai-projects/jet-long)  \nAbstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces routinely push the input an order of magnitude past the pretraining window, making zero-shot context extension the dominant deployment path for open-weight checkpoints. The dominant zero-shot methods (YaRN, Self-Extend, DCA) [1, 2 , 3] fix a single rescaling factor up front, so an aggressive factor sacrifices short-context fidelity while a conservative one breaks down at long contexts; recent length-aware variants [4, 5] adapt the mapping, but with a fitted or distance-dependent schedule. We propose Jet-Long, a tuning-free zero-shot method that pairs a local RoPE-faithful window with a long-range window whose rescaling factor adapts dynamically to the current sequence length via a parameter-free analytic schedule, recovering the base model exactly at short inputs while extrapolating cleanly at long ones. An inclusion–exclusion attention merge and an on-the-fly RoPE correction rotation make the bifocal construction essentially free at inference; fused into a single CuTe kernel, long-context prefill reaches up to 1.39 × FA2 throughput on H100 (approaching the Hopper-only FA4), and single-batch generation incurs ≤ 4% overhead at every length. On Qwen3- 1.7B/4B/8B [6] up to 128K context, Jet-Long leads RULER by +4 . 79/+2 . 18/+2 .03 pp over the strongest baseline at 1.7B/4B/8B, achieves the best overall accuracy on HELMET-RAG (a benchmark identified by HELMET as the most efficient predictor of downstream long-context performance [7]) and attains the lowest PG-19 perplexity. Jet-Long also generalizes to hybrid attention architectures such as Jet-Nemotron [8] for further long-context improvement without retraining, and remains hyperparameter-resilient for ease of deployment.  \nFigure 1 | Comparison between Jet-Long and baseline methods, applied on Qwen3-1.7B-Base, on per-length accuracy aggregated over all 13 RULER tasks and per-length perplexity in PG-19, alongside single-batch generation throughput on H100 at 128K context (the worst length we test) . Jet-Long achieves the highest accuracy and lowest perplexity at extended context lengths, preserves the base model’s pretrained performance within the training context, and incurs ≤ 4% latency overhead relative to FlashAttention-2 [9] .  \n1. Introduction  \nLarge language models (LLMs) are now deployed in long-context applications including long-document QA, repository-level code understanding, retrieval-augmented generation, and multi-step agentic workflows [10, 11 , 12 , 13 , 14 , 15 , 16 , 17 , 18] . The pressure is most severe in agentic LLMs that interleave reasoning, planning, and  \nCorresponding author(s): Han Cai ([hcai@nvidia.com](hcai@nvidia.com)).© 2026 NVIDIA. All rights reserved.  \ntool use across many turns [19, 20 , 21], and in coding agents operating over real software repositories [22, 23], where source code, execution traces, and tool outputs routinely accumulate to 100K+ tokens per task.  \nTraining directly at long context remains expensive: efficient kernels like FlashAttention [24] and Ring Attention [25] make memory linear but compute stays quadratic in sequence length, and long-context data is scarce while long-context fine-tuning often degrades short-context behavior [26, 27] . Models are therefore pretrained at a moderate window (4K–32K tokens) and expected to handle longer inputs at inference, a setting known as context extension [28, 29] . Zero-shot context extension (without fine-tuning) has become the dominant deployment mode for open-weight LLMs [6, 30 , 31], since a single release","cbCaivU0Q86ouMmp","https://ap.wps.com/l/cbCaivU0Q86ouMmp","pdf",1240466,9,1,14,"English","en",105,"# Introduction\n## Background and Motivation\n## Proposed Jet-Long Method\n## Contributions and Evaluation","[{\"question\":\"Why is zero-shot context extension important for open-weight LLM deployments?\",\"answer\":\"Because one released checkpoint must handle arbitrary downstream context lengths without fine-tuning. This makes zero-shot context extension the dominant deployment path for open-weight models.\"},{\"question\":\"What problem do existing zero-shot methods face when extending context length?\",\"answer\":\"They typically use a single rescaling factor or a fitted/distance-dependent schedule, which forces a trade-off: aggressive factors hurt short-context fidelity, while conservative ones break down at long contexts.\"},{\"question\":\"How does Jet-Long maintain fidelity at short inputs while extrapolating to long contexts?\",\"answer\":\"Jet-Long uses a bifocal construction with a local RoPE-faithful window and a long-range window whose rescaling factor adapts dynamically via a parameter-free analytic schedule, reproducing the base model exactly within its native context and extrapolating cleanly beyond it.\"}]",1784177590,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"jet-long-efficient-long-context-extension-with-dynamic-bifocal-rope","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/jet-long-efficient-long-context-extension-with-dynamic-bifocal-rope/82016/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is zero-shot context extension important for open-weight LLM deployments?","Question",{"text":76,"@type":77},"Because one released checkpoint must handle arbitrary downstream context lengths without fine-tuning. This makes zero-shot context extension the dominant deployment path for open-weight models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What problem do existing zero-shot methods face when extending context length?",{"text":81,"@type":77},"They typically use a single rescaling factor or a fitted/distance-dependent schedule, which forces a trade-off: aggressive factors hurt short-context fidelity, while conservative ones break down at long contexts.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Jet-Long maintain fidelity at short inputs while extrapolating to long contexts?",{"text":85,"@type":77},"Jet-Long uses a bifocal construction with a local RoPE-faithful window and a long-range window whose rescaling factor adapts dynamically via a parameter-free analytic schedule, reproducing the base model exactly within its native context and extrapolating cleanly beyond it.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]