[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85908-en":3,"doc-seo-85908-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85908,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","BOCCHI: A More Realistic and Challenging Benchmark for Local Motion Blur Detection with MSDCT-UNet","Local motion blur detection demands pixel-level localization of blurred regions, yet existing benchmarks can enable models to exploit gradient shortcuts that do not generalize. The work introduces BOCCHI, a real-captured, human-annotated local blur benchmark where sharp regions overlap the blur-gradient distribution, suppressing shortcut learning. It further proposes MSDCT-UNet, a frequency-aware encoder-decoder injecting multi-scale DCT priors via DCT Attention and FiLM. MSDCT-UNet achieves top in-domain mIoU and boundary localization and improves cross-dataset transfer with only 633 training images.","arXiv :2607 . 10427v1 [ cs .CV] 11 Jul 2026  \nBOCCHI: A More Realistic and Challenging Benchmark for Local Motion Blur Detection with  \nMSDCT-UNet  \nKuan-Lin Chen 1 , Yuan-Kang Lee2 , Cheng-Yuan Chiang 1 , and Jian-Jiun Ding 1  \n1 Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan  \n2 MediaTek Inc. , Hsinchu, Taiwan  \nAbstract. Local motion blur detection requires pixel-level localization of blurred regions. Existing benchmarks let models rely on gradient shortcuts that fail to transfer. We introduce BOCCHI (Blurred Objects Captured across Cameras with Human-annotated Imagery), a real-captured benchmark whose sharp regions overlap the blur gradient distribution and defeat these shortcuts, and propose MSDCT-UNet (Multi-Scale Discrete Cosine Transform UNet), a frequency-aware encoder-decoder injecting multi-scale DCT priors through DCT Attention and FiLM.  \nMSDCT-UNet ranks first in in-domain mIoU and boundary localization on BOCCHI, and BOCCHI-trained models outperform every other training source on cross-dataset transfer with only 633 training images.  \nKeywords: Motion blur detection · Local motion blur dataset · DCT attention · Frequency-domain learning · Semantic segmentation  \n1 Introduction  \nWhen an object moves during camera exposure, the resulting image exhibits spatially non-uniform local motion blur, where only moving objects appear blurred while the background remains sharp. Localizing such partial blur at pixel-level precision is challenging, and an accurate blur mask supports applications such as selective deblurring [45, 18], confidence calibration in object detection under motion, and restoration preprocessing. We formulate this problem as binary semantic segmentation.  \nExisting methods (surveyed in Sec. 2) mostly rely on hand-crafted frequency descriptors [34,31 ,6 ,23] or spatial-only learning [33,7 ,45 ,35 , 10 ,22], neither adapting well to diverse motion-blur scenes. We argue this arises partly from benchmark limitations: CUHKmotion [31] has only around 200 images, while ReLoBlur [17] concentrates on street scenes where the blurred regions are typically small (less than 30% of the image) and sit against textured backgrounds. Models trained on these benchmarks may learn shortcuts, e.g., treating low-frequency regions as blurred, without modeling the underlying optics of motion blur.  \nTo probe this harder regime, we introduce a real-captured benchmark whose sharp regions overlap the blur-region gradient distribution. As quantified in  \n2 K.-L. Chen et al.  \nBOCCHI Dataset  \nFig. 1: Overview of our contributions. Existing benchmarks can permit gradient-based shortcuts when blur and sharp regions are easily separable. BOCCHI addresses this with 633 real-captured, pixel-annotated images whose sharp regions cover both textured and smooth surfaces, creating strong gradient overlap with blurred objects. Built on this benchmark, MSDCT-UNet injects multi-scale DCT priors and achieves the best in-domain performance on BOCCHI, while BOCCHI-trained models show the strongest cross-dataset transfer with fewer training samples.  \nSec. 3 , this overlap is larger on BOCCHI than on any other training source, breaking the low-gradient-equals-blur shortcut and requiring models to reason about frequency cues rather than spatial gradient alone. The main contributions of this paper are:  \n1. BOCCHI benchmark. We introduce a real-captured, pixel-annotated dataset of 633 camera images exposing a harder local blur regime in which sharp regions span both textured backgrounds and smooth surfaces, creating gradientdistribution overlap with the blur region that is largely absent from prior datasets.  \n2. MSDCT-UNet. To our knowledge, the first segmentation network that injects multi-scale DCT priors at every encoder-decoder stage via a multi-head DCT Attention, FiLM modulation, and an AFASPP bottleneck, explicitly disentangling blur evidence from spatial appearance.  \n3. Empirical study. Against","cbCail95R4ljLrUr","https://ap.wps.com/l/cbCail95R4ljLrUr","pdf",12890617,5,1,28,"English","en",105,"# Introduction\n# Related Work\n## Blur Detection\n## Dense Prediction Architectures","[{\"question\":\"What problem does the paper target in local motion blur detection?\",\"answer\":\"It focuses on pixel-level localization of spatially non-uniform local motion blur, where moving objects are blurred while backgrounds remain sharp. Accurate blur masks are important for downstream tasks such as selective deblurring and restoration preprocessing.\"},{\"question\":\"What is BOCCHI and why is it considered more challenging than prior benchmarks?\",\"answer\":\"BOCCHI is a real-captured, pixel-annotated dataset of 633 camera images in a harder blur regime. Sharp regions overlap the blur-region gradient distribution, which breaks gradient-based shortcuts available in earlier benchmarks.\"},{\"question\":\"How does MSDCT-UNet improve performance compared with standard approaches?\",\"answer\":\"MSDCT-UNet is a frequency-aware encoder-decoder that injects multi-scale DCT priors at every stage. Using DCT Attention and FiLM modulation, it encourages the model to reason about frequency cues rather than relying on spatial gradient alone.\"}]",1784207101,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"bocchi-a-more-realistic-and-challenging-benchmark-for-local-motion-blur-detection-with-msdct-unet","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/bocchi-a-more-realistic-and-challenging-benchmark-for-local-motion-blur-detection-with-msdct-unet/85908/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper target in local motion blur detection?","Question",{"text":76,"@type":77},"It focuses on pixel-level localization of spatially non-uniform local motion blur, where moving objects are blurred while backgrounds remain sharp. Accurate blur masks are important for downstream tasks such as selective deblurring and restoration preprocessing.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is BOCCHI and why is it considered more challenging than prior benchmarks?",{"text":81,"@type":77},"BOCCHI is a real-captured, pixel-annotated dataset of 633 camera images in a harder blur regime. Sharp regions overlap the blur-region gradient distribution, which breaks gradient-based shortcuts available in earlier benchmarks.",{"name":83,"@type":74,"acceptedAnswer":84},"How does MSDCT-UNet improve performance compared with standard approaches?",{"text":85,"@type":77},"MSDCT-UNet is a frequency-aware encoder-decoder that injects multi-scale DCT priors at every stage. Using DCT Attention and FiLM modulation, it encourages the model to reason about frequency cues rather than relying on spatial gradient alone.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]