[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81600-en":3,"doc-seo-81600-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81600,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking","Efficient long-sequence processing with Transformer models often relies on context parallelism across accelerators, yet common methods primarily scale compute over the context dimension and still run into activation-memory limits. Techniques like pipelined distributed Transformers or activation offloading extend context length at the cost of training throughput. UPipe introduces fine-grained, attention head–level chunking that reduces self-attention activation memory by up to 87.5% for 32B models while matching prior training speed, enabling 5M-token training on Llama3-8B using a single 8×H100 node.","Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking  \nRavi Ghadia 1 Maksim Abraham 1 Sergei Vorobyov 1 Max Ryabinin 1  \narXiv :2602 .2 1 196v2 [ cs .LG] 10 Jul 2026  \nAbstract  \nEfficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism. The dominant approaches in this family of methods, such as Ring Attention or DeepSpeed Ulysses, enable scaling over the context dimension but do not focus on memory efficiency, which limits the sequence lengths they can support. More advanced techniques, such as Fully Pipelined Distributed Transformer or activation offloading, can further extend the possible context length at the cost of training throughput. In this paper, we present UPipe, a simple yet effective context parallelism technique that performs fine-grained chunking at the attention head level.  \nThis technique significantly reduces the activation memory usage of self-attention, breaking the activation memory barrier and unlocking much longer context lengths. Our approach reduces intermediate tensor memory usage in the attention layer by as much as 87.5% for 32B Transformers, while matching previous context parallelism techniques in training speed. UPipe can support the context length of 5M tokens when training Llama3-8Bon a single 8×H100 node, improving upon prior methods by over 25% .  \n1. Introduction  \nThe Transformer architecture (Vaswani et al., 2017) has powered significant advances in AI in recent years, ranging from language models with agentic and reasoning capabilities (Gemini Team, 2025 ; Kimi Team et al., 2025 ; Chenet al., 2025) to video generation (Team Wan et al., 2025 ; Wu et al., 2025 ; [Sand.ai](Sand.ai) et al., 2025) . As the field continues to progress, the demand for longer context lengths in AI models continues to grow due to applications such as code  \n1Together AI. Correspondence to: Ravi Ghadia \u003Crgha[dia@utexas.edu](dia@utexas.edu) >, Max Ryabinin \u003C[mryab@together.ai](mryab@together.ai) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nThroughput (TPS)  \n500  \n400  \n300  \n200  \n100  \n0  \n1M 2M 3M 4M 5M  \nSequence length  \nFigure 1. Comparison of context parallelism approaches on longsequence training for Llama3-8B using 8 × H100s. UPipe provides maximum efficiency, resulting in a longer maximum context length (5M tokens) while retaining throughput.  \ngeneration (Li et al., 2023a ; Hui et al., 2024), long document understanding (Chia et al., 2024 ; Jiang et al., 2024), or even audio processing (Hori et al., 2021) . However, training models to effectively process such long sequences is limited by the accelerator hardware: beyond a certain limit, even keeping the activations necessary for self-attention becomes a bottleneck. As a result, methods that reduce the memory requirements of long-context training have recently attracted a surge of research interest.  \nThe most scalable approaches for increasing the context size beyond a single accelerator leverage distributed training, splitting the computations and memory allocations across multiple devices. In particular, the context parallelism (Liet al., 2022 ; 2023b ; Jacobs et al., 2023) family of methods (also known as sequence parallelism) focuses on sharding model operations across the sequence axis. These methods enable effective scaling in context length with the number of accelerators, but the activation memory per device still scales linearly with the sequence length. Therefore, at very long sequence lengths (>2M), the activation memory starts to become a bottleneck, limiting the training capacity.  \nIn this paper, we propose UPipe, a context parallelism method that focuses on improving the memory efficiency of long-context training while maintaining performance on par with current approaches. Our method is designed on the principle that for lon","cbCaivMntkKHlHhQ","https://ap.wps.com/l/cbCaivMntkKHlHhQ","pdf",891344,4,1,16,"English","en",105,"# Abstract\n# Introduction\n## Motivation: long-context limits and activation memory bottlenecks\n## UPipe: headwise chunking for improved memory reuse\n## Results and throughput on Llama 3-8B and larger models\n# Contributions","[{\"question\":\"What problem does UPipe address in long-context Transformer training?\",\"answer\":\"UPipe targets the activation-memory bottleneck that appears at very long sequence lengths, which limits how large the context can be during training.\"},{\"question\":\"How does UPipe reduce activation memory in the attention layer?\",\"answer\":\"UPipe performs fine-grained chunking at the attention head level, serializing attention execution in multiple stages so intermediate tensors can be reused more effectively.\"},{\"question\":\"What context lengths and hardware setups does UPipe enable?\",\"answer\":\"On Llama 3-8B, UPipe supports up to 5 million tokens when training on a single 8×H100 node, and up to 8M tokens on two H100 nodes with unified sequence parallelism, with throughput comparable to existing methods.\"}]",1784174632,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"untied-ulysses-memory-efficient-context-parallelism-via-headwise-chunking","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/untied-ulysses-memory-efficient-context-parallelism-via-headwise-chunking/81600/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does UPipe address in long-context Transformer training?","Question",{"text":75,"@type":76},"UPipe targets the activation-memory bottleneck that appears at very long sequence lengths, which limits how large the context can be during training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does UPipe reduce activation memory in the attention layer?",{"text":80,"@type":76},"UPipe performs fine-grained chunking at the attention head level, serializing attention execution in multiple stages so intermediate tensors can be reused more effectively.",{"name":82,"@type":73,"acceptedAnswer":83},"What context lengths and hardware setups does UPipe enable?",{"text":84,"@type":76},"On Llama 3-8B, UPipe supports up to 5 million tokens when training on a single 8×H100 node, and up to 8M tokens on two H100 nodes with unified sequence parallelism, with throughput comparable to existing methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]