[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-146482-105":59,"doc-detail-146482-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","srcb-at-semeval-2022-task-5-pretraining-based-image-to-text-late-sequential-fusion-system-for-multimodal-misogynous-meme-identification","SRCB at SemEval-2022 Task 5 - Pretraining Based Image to Text Late Sequential Fusion System for Multimodal Misogynous Meme Identification","","Online misogyny meme detection is a multimodal classification problem where the relationship between image and text strongly affects fusion learning. This work evaluates single-stream UNITER and dual-stream CLIP multimodal pretrained models for handling both strongly and weakly correlated image-text pairs. An XGBoost classifier using CLIP image features achieves the best overall performance and remains robust under domain shift. Building on this, the proposed PBR system ensembles pretrained models, boosting, and rule-based adjustment, integrating text through a late sequential fusion scheme.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/srcb-at-semeval-2022-task-5-pretraining-based-image-to-text-late-sequential-fusion-system-for-multimodal-misogynous-meme-identification/146482/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/srcb-at-semeval-2022-task-5-pretraining-based-image-to-text-late-sequential-fusion-system-for-multimodal-misogynous-meme-identification/146482.png","ImageObject",300,407,{"name":92,"@type":93},"Aria","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-16","2026-08-26",true,{"@type":102,"interactionType":103,"userInteractionCount":81},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What is the main goal of SemEval-2022 Task 5 Multimedia Automatic Misogyny Identification?","Question",{"text":112,"@type":113},"It focuses on identifying misogynous memes using both the meme image and its overlaying text, with two sub-tasks for coarse (misogynous vs. not) and fine-grained category classification.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"How do the approaches differ between UNITER and CLIP in this work?",{"text":117,"@type":113},"UNITER is used as a single-stream model that fuses image and text early, while CLIP is treated as a dual-stream model with separate image/text encoders for cross-modal learning.",{"name":119,"@type":110,"acceptedAnswer":120},"Why is an ensemble late sequential fusion scheme used in the proposed PBR system?",{"text":121,"@type":113},"Because the method aims to combine pretrained models, boosting, and rule-based adjustment while fusing text into the classification using a late sequential strategy to better handle varying image-text correlation.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},146482,1787743568,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":81,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":41},2336464648322,"https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488","SRCB at SemEval-2022 Task 5: Pretraining Based Image to Text Late Sequential Fusion System for Multimodal Misogynous Meme Identification  \nYujin WangJing ZhangB, Bohua PengXudong Zhang∗  \nXiaoyan Qu, Yimeng Zhuang, Song Liu  \nSamsung Research China-Beijing (SRC-B)  \n{[yujin1.wang](yujin1.wang) , jing97.zhang, [bohua.peng}@samsung.com](bohua.peng}@samsung.com)[ ](bohua.peng}@samsung.com){xudong.zhang, [xiaoyan11.qu}@samsung.com](xiaoyan11.qu}@samsung.com)[ ](xiaoyan11.qu}@samsung.com){ym.zhuang, [s0101.liu}@samsung.com](s0101.liu}@samsung.com)  \nAbstract  \nOnline misogyny meme detection is an image/text multimodal classification task, the complicated relation of image and text challenges the intelligent system’s modality fusion learning capability. In this paper, we investigate the single-stream UNITER and dual-stream CLIP multimodal pretrained models on their capability to handle strong and weakly correlated image/text pairs. The XGBoost classifier with image features extracted by the CLIP model has the highest performance and being robust on domain shift. Based on this, we propose the PBR system, an ensemble system of Pretraining models, Boosting method and Rule-based adjustment, text information is fused into the system using our late sequential fusion scheme.  \nOur system ranks 1st place on both sub-task  \nA and sub-task B of the SemEval-2022 Task  \n5 Multimedia Automatic Misogyny Identification, with 0 . 834/0 .731 macro F1 scores for subtask A/B correspondingly.  \n1 Introduction  \nMuch of the real world’s information comes in multimodality, a combination of images, texts, audiosand so on. Multimodal understanding aims to utilize different modal of information to improve the overall system recognition intelligence or robustness (Gadzicki et al., 2020), which plays a key foundation role in cognitive AI and embodied AI.  \nWith transfer learning by large deep models and colossal corpus achieving remarkable success in vision and language domain, there is a rising interest in combining both sides’ advances to push the multimodality understanding further (Lu et al., 2019 ; Tan and Bansal, 2019 ; Chen et al., 2019 ; Li et al., 2020 ; Yu et al., 2020 ; Huo et al., 2021 ; Kim et al., 2021 ; Radford et al., 2021) . We will limit the discussion scope of multimodal to vision and language in this paper. There are two kinds of representative  \nontribution during Intership in Samsung Research China-Beijing.  \narchitecture of multimodal learning models, singlestream models and dual-stream models. Singlestream model fuses the image and text data at an early stage, and then feed into the model. Dualstream models design separated structure as image encoder and text encoder, and a further module is stacked on top of the unimodel encoders for crossmodal learning objectives (Tan and Bansal, 2019 ; Yu et al., 2020 ; Radford et al., 2021 ; Huo et al., 2021) . Usually per-unimodal objectives and multimodal objectives are designed to ensure that the model learns unimodal and crossmodal knowledge, like masked image prediction, masked token prediction, and text-image pairing (Chen et al., 2019 ; Kim et al., 2021) . Two kinds of data distributions are explored for the large-scale pretraining, strongly paired data (Chen et al., 2019 ; Radford et al., 2021 ; Li et al., 2020 ; Kim et al., 2021) and weakly paired data (Huo et al., 2021) . The different distributions would directly affect the correlations learned by the model, yet each pretraining corpus only falls in one pattern.  \nThe SemEval-2022 Task 5 (Fersini et al., 2022) Multimedia Automatic Misogyny Identification (MAMI) is a multimodal classification task in English. It targets the identification of misogynous memes (characterized by a pictorial content with an overlaying text a posteriori introduced by human), using the image and text from the meme as input data. It has two sub-tasks: sub-task A: 2-fold classification, to identify whether a meme is misogynous or not; sub-task B: 4-fold fine-grai","cbCairxQp1wwfit8","https://ap.wps.com/l/cbCairxQp1wwfit8","pdf",4944052,12,"English","# Abstract\n# Introduction\n## Multimodal learning background\n## SemEval-2022 Task 5 and sub-tasks","[{\"question\":\"What is the main goal of SemEval-2022 Task 5 Multimedia Automatic Misogyny Identification?\",\"answer\":\"It focuses on identifying misogynous memes using both the meme image and its overlaying text, with two sub-tasks for coarse (misogynous vs. not) and fine-grained category classification.\"},{\"question\":\"How do the approaches differ between UNITER and CLIP in this work?\",\"answer\":\"UNITER is used as a single-stream model that fuses image and text early, while CLIP is treated as a dual-stream model with separate image/text encoders for cross-modal learning.\"},{\"question\":\"Why is an ensemble late sequential fusion scheme used in the proposed PBR system?\",\"answer\":\"Because the method aims to combine pretrained models, boosting, and rule-based adjustment while fusing text into the classification using a late sequential strategy to better handle varying image-text correlation.\"}]","SRCB at SemEval-2022 Task 5 - Pretraining Based Image to Text Late Sequential Fusion System for Multimodal Misogynous Meme Identification | PDF"]