[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86137-en":3,"doc-seo-86137-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86137,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder","U-Net-style encoder-decoder models rely on accurate top-down decoding to reconstruct low-level details from high-level semantics, which requires effective multi-scale feature fusion. Existing attention-based fusion typically computes attention from global decoder features or from global–local correlation, then modulates encoder features. This work proposes difference-driven gating that derives attention weights from the difference between the two feature streams: Feature-difference gating (FDG) and Entropy-difference gating (EDG). Coupled gating maps modulate both streams and outperform prior attention methods, with EDG achieving the strongest results.","Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder  \nKai Li, Student Member, IEEE, Xuechao Zou, Jiashen Fu, Zijun Yan, Xintong Wang, and Xiaolin  \nHu, Senior Member, IEEE  \narXiv :2607 . 11096v1 [ cs .CV] 13 Jul 2026  \nAbstract—The U-Net style models have been widely used in many applications. A critical step in these models is to reconstruct the lower-level features using a top-down decoder. This reconstruction requires precise fusion of high-level semantics and low-level details. Existing attention-based fusion methods typically derive attention weights from the top-down decoder features (global) alone or the correlation between the top-down decoder features and the bottom-up encoder features (local), then modulate the encoder features using these weights. In this work, we explore a different paradigm: deriving attention weights from the difference between the two feature streams. To this end, we propose two difference-based gating approaches: Feature-difference gating (FDG), which directly uses the absolute difference between global and local features to generate adaptive gating maps, and Entropy-difference gating (EDG), which measures the representational certainty of each stream via information entropy and uses their signed entropy difference to derive the attention weights. Both methods produce coupled gating maps that simultaneously modulate the global and local features. Experiments on different tasks including medical image segmentation, remote sensing image cloud removal and speech separation showed that both methods outperformed existing attention-based fusion methods, and EDG performed better. The results suggested a new paradigm for multi-scale feature fusion in the U-Net style structures.  \nIndex Terms—Entropy-guided feature fusion, certainty-aware gating, U-Net-based architectures, medical image segmentation, cloud removal, speech separation.  \n~~ ~~ ✦ ~~ ~~  \n1 INTRODUCTION  \nTHE U-Net architecture [1], a classical encoder-decoder  \nframework, is widely used in various dense prediction tasks, including image processing [2], [3], [4], [5], [6] and audio processing [7], [8] . This architecture processes information in a bottom-up then top-down manner, analogous to that in biological visual systems [9], [10] . It comprises a contracting encoder that abstracts low-level details into highlevel semantic representations and an expanding decoder that leverages skip connections to fuse preserved spatial details with the upsampled semantic context [11] .  \nIn the U-Net-based architecture, the effectiveness of the decoder relies on how well it integrates multi-scale information. Therefore, the fusion mechanism that integrates the fine-grained details from the encoder with the semantic context from the decoder becomes a critical bottleneck for performance [12], [13] . Within the U-Net architecture shown in Fig. 1, this generic fusion mechanism is instantiated at each decoding stage by a fusion module. Existing fusion methods can be broadly divided into two categories. The first category uses simple fusion mechanisms, which mainly rely on element-wise addition or concatenation [1], [14],[15], as shown in Fig. 2(a) and (b) . Although these meth-  \n• K. Li and X. Hu are with the Department of Computer Science and Technology, Institute for Artificial Intelligence, BNRist, IDG/McGovern Institute for Brain Research, Tsinghua University, Beijing, China. X. Hu is also with the Chinese Institute for Brain Research (CIBR), Beijing, China.  \n• X. Zou is with the School of Computer Science and Technology, Beijing Jiaotong University, Beijing, China.  \n• J. Fu, Z. Yan and X. Wang are with the Department of Computer Science and Technology, Tsinghua University, Beijing, China.  \n(Corresponding authors: Xiaolin Hu.)  \nFig. 1. Overview of U-Net architecture with integrated fusion modules. Following the neuroscience convention, the U-Net is depicted in an inverted manner, with the coarsest features placed at the top. Within","cbCaikqeKZzbFfyw","https://ap.wps.com/l/cbCaikqeKZzbFfyw","pdf",14281429,4,1,15,"English","en",105,"# Introduction\n## Decoder fusion as a performance bottleneck\n## Existing fusion methods: simple vs. attention-based\n## Motivation for difference-driven attention","[{\"question\":\"What fusion problem does the work address in U-Net decoders?\",\"answer\":\"It targets how to integrate multi-scale information during decoding by reconstructing low-level details using high-level semantics, where the fusion mechanism at each decoding stage can limit overall performance.\"},{\"question\":\"How do the proposed methods compute attention weights?\",\"answer\":\"They derive attention from the difference between global (decoder) and local (encoder) feature streams: FDG uses the absolute feature difference, while EDG uses signed differences in information entropy to reflect representational certainty.\"},{\"question\":\"What tasks are used to evaluate the gating approaches, and which method performs best?\",\"answer\":\"Experiments include medical image segmentation, remote sensing image cloud removal, and speech separation. Both FDG and EDG outperform existing attention-based fusion methods, and EDG delivers better results overall.\"}]",1784208846,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"difference-driven-gating-adaptive-feature-fusion-for-u-net-decoder","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/difference-driven-gating-adaptive-feature-fusion-for-u-net-decoder/86137/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What fusion problem does the work address in U-Net decoders?","Question",{"text":75,"@type":76},"It targets how to integrate multi-scale information during decoding by reconstructing low-level details using high-level semantics, where the fusion mechanism at each decoding stage can limit overall performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the proposed methods compute attention weights?",{"text":80,"@type":76},"They derive attention from the difference between global (decoder) and local (encoder) feature streams: FDG uses the absolute feature difference, while EDG uses signed differences in information entropy to reflect representational certainty.",{"name":82,"@type":73,"acceptedAnswer":83},"What tasks are used to evaluate the gating approaches, and which method performs best?",{"text":84,"@type":76},"Experiments include medical image segmentation, remote sensing image cloud removal, and speech separation. Both FDG and EDG outperform existing attention-based fusion methods, and EDG delivers better results overall.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]