[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81831-en":3,"doc-seo-81831-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":20},81831,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","3DLS: 3D Logic-Stacked Architecture for Disaggregated LLM Serving","Large language model (LLM) serving increasingly uses prefill-decode (PD) disaggregation with tensor parallelism (TP) to handle large models and long contexts. In conventional 2D/2.5D chiplet designs, layer-wise KV-cache transfers and decode-side TP collectives contend for the same lateral die-to-die (D2D) interconnect, increasing communication latency and stretching token generation intervals. 3DLS proposes logic-on-logic 3D-stacked routing that isolates KV-cache traffic via vertical interconnects while keeping TP collectives on the lateral fabric, improving throughput up to 1.49× and reducing end-to-end latency by up to 60.2%.","3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving  \nJaehun Lee, In-Jun Jung, and Joo-Young Kim, Senior Member, IEEE  \narXiv :2607 .0 16 17v 1 [ cs .AR] 2 Jul 2026  \nAbstract—Large language model (LLM) serving increasingly combines prefill-decode (PD) disaggregation with tensor parallelism (TP) to support large models and long contexts. In conventional 2D/2.5D chiplet architectures, layer-wise prefill-todecode KV-cache transfer decode-side TP collectives share the same lateral die-to-die (D2D) interconnect, creating mixed-traffic contention on the decode critical path. This contention increases communication latency, prolongs token generation intervals, and degrades end-to-end serving performance. We propose 3DLS, a logic-on-logic 3D-stacked chiplet architecture that separates traffic classes by routing KV-cache transfers through vertical interconnects while preserving decode-side TP collectives on the lateral D2D fabric. 3DLS achieves up to 1.49× throughput and 60.2% lower end-to-end (E2E) latency over the sharedfabric planar baseline, and still achieves up to 1.17× throughput and 31.4% lower E2E latency over a workload-aware prioritymanaged planar baseline. These results highlight that physical isolation is an important design principle for future chiplet-based PD-disaggregated LLM serving systems.  \nIndex Terms—Large Language Model, LLM Serving, Disaggregated Serving, Chiplet, KV cache, 3D integration.  \nI. INTRODUCTION  \nLARGE language model (LLM) serving is shifting toward  \nlarger models and longer context windows, increasing compute and memory resource demands. At the same time, satisfying these demands with monolithic large-die accelerators has become increasingly difficult due to the rising cost of advanced process nodes, yield challenges, and reticle-size limits. These trends make chiplet-based integration a promising solution for scaling LLM inference hardware beyond the practical limits of monolithic designs [1], making on-package communication a major concern for serving performance.  \nModern LLM inference consists of two phases with distinct characteristics: a compute-intensive prefill phase and a memory-bound decode phase. Because these phases differ in resource demands, recent serving systems increasingly adopt prefill-decode (PD) disaggregation [2], [3], assigning each phase to a dedicated resource pool to improve efficiency and utilization. Meanwhile, serving large models continues to rely on tensor parallelism (TP) [4] to meet memory-capacity and throughput requirements. However, TP introduces frequent  \n© 2026 IEEE. This is the author’s accepted manuscript of an article accepted for publication in IEEE Computer Architecture Letters. DOI: 10.1109/LCA.2026.3709108 . Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses.  \nJaehun Lee is with the Graduate School of System Architect, KAIST, Daejeon 34141, South Korea (e-mail: [jaehunlee@kaist.ac.kr](jaehunlee@kaist.ac.kr)).  \nIn-Jun Jung and Joo-Young Kim are with the School of Electrical Engineering, KAIST, Daejeon 34141, South Korea (e-mail: [injun@kaist.ac.kr](injun@kaist.ac.kr); [jooyoung1203@kaist.ac.kr](jooyoung1203@kaist.ac.kr)).  \n(a) TP=4 case, 2 hop (b) TP=16 case, 4 hop  \n :Prefill die Collective Communication  :Decode die  : KV cache transfer  \nFig. 1. Traffic contention in conventional 2D/2.5D chiplet-based PDdisaggregated serving. (a) TP=4 case with 2-hop KV cache transfer; (b) TP=16 case with 4-hop KV cache transfer.  \ncollective operations whose overhead grows with model scale and TP degree. As a result, a TP-enabled PD-disaggregated serving system must simultaneously support two heterogeneous traffic classes: KV-cache transfer from prefill to decode, and latency-critical decode-side collective communication.  \nFig. 1 shows that when PD-disaggregated serving with TP is mapped onto a conventional 2D/2.5D chiplet architecture [5], KV-cache transfer and decode-side TP collectives are fo","cbCaijTdbWsXOwBL","https://ap.wps.com/l/cbCaijTdbWsXOwBL","pdf",779033,10,1,4,"English","en",105,"# Introduction\n# Traffic Contention in Conventional 2D/2.5D Chiplets\n# Proposed 3DLS Architecture\n# Communication Operations and Conflict Analysis","[{\"question\":\"What problem does 3DLS address in PD-disaggregated LLM serving with tensor parallelism?\",\"answer\":\"It addresses contention on the lateral die-to-die interconnect between KV-cache transfers (prefill to decode) and decode-side TP collective communication, which increases latency and degrades end-to-end performance.\"},{\"question\":\"How does 3DLS physically isolate the two traffic classes?\",\"answer\":\"3DLS routes KV-cache transfers through dedicated vertical 3D interconnects, while keeping decode-side TP collectives on the lateral D2D fabric.\"},{\"question\":\"What performance improvements does the document report for 3DLS?\",\"answer\":\"It reports up to 1.49× throughput and 60.2% lower end-to-end latency versus a shared-fabric planar baseline, and up to 1.17× throughput with 31.4% lower end-to-end latency versus a workload-aware priority-managed planar baseline.\"}]","3DLS: 3D Logic-Stacked Architecture for Disaggregated LLM Serving | PDF",1784176464,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":29},"3dls-3d-logic-stacked-architecture-for-disaggregated-llm-serving","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":22},"https://docshare.wps.com/document/3dls-3d-logic-stacked-architecture-for-disaggregated-llm-serving/81831/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-30","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does 3DLS address in PD-disaggregated LLM serving with tensor parallelism?","Question",{"text":75,"@type":76},"It addresses contention on the lateral die-to-die interconnect between KV-cache transfers (prefill to decode) and decode-side TP collective communication, which increases latency and degrades end-to-end performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does 3DLS physically isolate the two traffic classes?",{"text":80,"@type":76},"3DLS routes KV-cache transfers through dedicated vertical 3D interconnects, while keeping decode-side TP collectives on the lateral D2D fabric.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance improvements does the document report for 3DLS?",{"text":84,"@type":76},"It reports up to 1.49× throughput and 60.2% lower end-to-end latency versus a shared-fabric planar baseline, and up to 1.17× throughput with 31.4% lower end-to-end latency versus a workload-aware priority-managed planar baseline.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":20,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]