[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86165-en":3,"doc-seo-86165-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86165,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","TC-MAF Train-Calibrated Bounded Multi-Evidence Fusion for Multimodal Industrial Anomaly Detection","Multimodal anomaly detection leverages complementary RGB and 3D evidence, but RGB reconstruction reliability varies across product categories and class-wise test-time evidence selection is often unavailable. TC-MAF introduces a base-anchored multi-evidence fusion framework that unifies multimodal detection, complementary Dinomaly evidence, and a cross-modal consistency cue under a fixed pixel-level fusion formula. A lightweight training-dispersion confidence term scales auxiliary participation using only normal statistics, yielding strong MVTec-3D performance: 0.979 AUROC image-level and 0.990 AUPRO pixel-level, with fusion structure as the dominant driver.","TC-MAF: Train-Calibrated Bounded Multi-Evidence Fusion for Multimodal Industrial Anomaly Detection  \nMing Deng  \nShanghai University Shanghai, China [chranos@shu.edu.cn](chranos@shu.edu.cn)  \nXiaochuan Hu  \nUniversity of Electronic Science and Technology of China Chengdu, China [202421090208@std.uestc.edu.cn](202421090208@std.uestc.edu.cn)  \nSijin Sun  \nNational University of Singapore Singapore, Singapore [sun.sijin@u.nus.edu](sun.sijin@u.nus.edu)  \nXing Wu∗ Shanghai University Shanghai, China [xingwu@shu.edu.cn](xingwu@shu.edu.cn)  \narXiv :2607 . 11170v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nMultimodal anomaly detection benefits from complementary RGB and 3D evidence, yet auxiliary RGB reconstruction is not equally reliable across product categories and class-wise test-time policy selection is usually unavailable. We propose TC-MAF, a baseanchored multi-evidence fusion design that combines a multimodal detector, complementary Dinomaly evidence, and a small crossmodal consistency cue under one fixed pixel-level fusion formula. A lightweight training-dispersion confidence (TDC) term scales auxiliary participation using only normal training statistics. On MVTec-3D, TC-MAF reaches 0.979 image-level AUROC and 0.990 pixel-level AUPRO, achieving the best mean results on both detection and localization among the compared multimodal methods. Systematic ablations show that the fusion structure itself is the dominant factor, while TDC provides a smaller but reproducible calibration gain over no calibration or arbitrary calibration. Additional experiments show that the same design remains effective under a pooled-statistics variant, auxiliary-branch and backbone substitutions, few-shot settings, a missing-3D setting, and cross-dataset evaluation on Eyecandies. Code is available at [https://anonymous.4open.science/r/TC_MAF-C3BB](https://anonymous.4open.science/r/TC_MAF-C3BB).  \nCCS Concepts  \n• Computing methodologies → Anomaly detection; Image representations; Computer vision tasks.  \nKeywords  \nmultimodal anomaly detection, industrial anomaly detection, RGBD anomaly detection, adaptive fusion, known-category anomaly detection  \n1 Introduction  \nCombining RGB appearance and 3D geometry improves industrial anomaly detection because many defects manifest more clearly in one modality than the other [1, 26, 28, 33] . Once multiple evidence sources are available, however, the integration rule becomes important: auxiliary evidence may help substantially on some product categories and much less on others, while anomalous validation data for category-wise policy selection is usually unavailable. Prior methods advance representation learning, anomaly generation, or  \n∗ Corresponding author.  \nFigure 1: Detection-localization operating points on MVTec- 3D. Each point shows one multimodal method in the IAUROC–P-AUPRO plane. TC-MAF occupies the upper-right frontier among the compared methods, with the highest mean I-AUROC and P-AUPRO.  \ncross-modal interaction [8, 16, 20, 25, 28], but usually aggregate branch outputs with fixed rules that are only lightly analyzed.  \nThis issue becomes especially visible when a multimodal detector is paired with an RGB reconstruction branch. Reconstruction can recover subtle appearance anomalies that geometry may miss, yet its reliability is category-dependent: identity shortcuts emerge under diverse normal patterns [30], and stable multi-class reconstruction requires careful design [13, 32] . The resulting problem is not simply to add more evidence, but to incorporate that evidence without allowing it to overwrite the stronger multimodal base indiscriminately.  \nWe address this with TC-MAF, a base-anchored multi-evidence fusion design. TC-MAF keeps the multimodal detector as the primary source, adds complementary Dinomaly evidence [13] through a constrained pixel-level mixture, and uses a small cross-modal consistency (CMC) term to reinforce regions where both modalities  \nappear inconsistent. A lightweight training-","cbCaibIGii5UxB9y","https://ap.wps.com/l/cbCaibIGii5UxB9y","pdf",8989997,3,1,13,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does TC-MAF address in multimodal industrial anomaly detection?\",\"answer\":\"TC-MAF targets the uneven reliability of auxiliary RGB reconstruction across product categories and the lack of available class-wise test-time policy selection for evidence integration.\"},{\"question\":\"How does TC-MAF fuse multiple evidence sources?\",\"answer\":\"TC-MAF keeps the multimodal detector as the anchor, incorporates complementary Dinomaly evidence via a constrained pixel-level mixture, and adds a cross-modal consistency cue to reinforce regions with cross-modality disagreement.\"},{\"question\":\"What is the role of the training-dispersion confidence (TDC) term?\",\"answer\":\"TDC scales the auxiliary participation using only normal training statistics, providing calibration that is smaller than the fusion-structure effect but still reproducible and beneficial.\"}]",1784209036,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tc-maf-train-calibrated-bounded-multi-evidence-fusion-for-multimodal-industrial-anomaly-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/tc-maf-train-calibrated-bounded-multi-evidence-fusion-for-multimodal-industrial-anomaly-detection/86165/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TC-MAF address in multimodal industrial anomaly detection?","Question",{"text":75,"@type":76},"TC-MAF targets the uneven reliability of auxiliary RGB reconstruction across product categories and the lack of available class-wise test-time policy selection for evidence integration.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TC-MAF fuse multiple evidence sources?",{"text":80,"@type":76},"TC-MAF keeps the multimodal detector as the anchor, incorporates complementary Dinomaly evidence via a constrained pixel-level mixture, and adds a cross-modal consistency cue to reinforce regions with cross-modality disagreement.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the training-dispersion confidence (TDC) term?",{"text":84,"@type":76},"TDC scales the auxiliary participation using only normal training statistics, providing calibration that is smaller than the fusion-structure effect but still reproducible and beneficial.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]