[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124807-en":3,"doc-seo-124807-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124807,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","EyeXNet - Enhancing Abnormality Detection and Diagnosis via Eye-Tracking and X-ray Fusion - Abstract","Integrating eye gaze data with chest X-ray images in deep learning has produced conflicting results in prior studies. EyeXNet addresses this by arguing that researchers often ignore the human element of eye tracking and apply raw signals without proper preprocessing. The work introduces EyeXNet, a multimodal model combining CXR images with radiologists’ fixation masks to predict abnormality locations. Fixation maps are analyzed around reporting moments, showing more targeted focus during reporting. Across eight experiments, fixation-mask integration improves recall and precision versus baseline, supporting human-centered multimodal learning for CXR analysis.","machine learning & knowledge extraction  \nArticle  \nEyeXNet: Enhancing Abnormality Detection and Diagnosis via Eye-Tracking and X-ray Fusion  \nChihcheng Hsieh 1,†, André Luís 2,3,†, José Neves 2,3, Isabel Blanco Nobre 4, Sandra Costa Sousa 4, Chun Ouyang 1, Joaquim Jorge 2,3 and Catarina Moreira 1,2,5, *  \nCitation: Hsieh, C.; Luís, A.; Neves, J.; Nobre, I.B.; Sousa, S.C.; Ouyang, C.; Jorge, J.; Moreira, C. EyeXNet: Enhancing Abnormality Detection and Diagnosis via Eye-Tracking and X-ray Fusion. Mach. Learn. Knowl. Extr. 2024, 6, 1055–1071. [https://](https://)[ ](https://)[doi.org/10.3390/make6020048](doi.org/10.3390/make6020048)  \nAcademic Editor: Andreas Holzinger  \nReceived: 24 March 2024  \nRevised: 19 April 2024  \nAccepted: 6 May 2024  \nPublished: 9 May 2024  \nCopyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://](https://)[ ](https://)[creativecommons.org/licenses/by/](creativecommons.org/licenses/by/)[ ](creativecommons.org/licenses/by/)[4.0/](4.0/)) .  \n1 School of Information Systems, Queensland University of Technology, Brisbane, QLD 4000, Australia; [chihcheng.hsieh@hdr.qut.edu.au](chihcheng.hsieh@hdr.qut.edu.au) (C.H.); [c.ouyang@qut.edu.au](c.ouyang@qut.edu.au) (C.O.)  \n2 Instituto Superior Técnico, Universidade de Lisboa, 1049-001 Lisboa, Portugal; [andre.t.luis@tecnico.ulisboa.pt](andre.t.luis@tecnico.ulisboa.pt) (A.L.); [jose.s.neves@tecnico.ulisboa.pt](jose.s.neves@tecnico.ulisboa.pt) (J.N.); [jorgej@tecnico.ulisboa.pt](jorgej@tecnico.ulisboa.pt) (J.J.)  \n3 INESC-ID, 1000-029 Lisbon, Portugal  \n4 Grupo Lusíadas, Imagiology Department, 1500-458 Lisbon, Portugal; [isabel.blanco.nobre@lusiadas.pt](isabel.blanco.nobre@lusiadas.pt) (I.B.N.); [sandra.costa.sousa@lusiadas.pt](sandra.costa.sousa@lusiadas.pt) (S.C.S.)  \n5 Human Technology Institute, University of Technology Sydney, Sydney, NSW 2007, Australia  \n* Correspondence: [catarina.pintomoreira@uts.edu.au](catarina.pintomoreira@uts.edu.au)[ ](catarina.pintomoreira@uts.edu.au)† These authors contributed equally to this work.  \nAbstract: Integrating eye gaze data with chest X-ray images in deep learning (DL) has led to contradictory conclusions in the literature. Some authors assert that eye gaze data can enhance prediction accuracy, while others consider eye tracking irrelevant for predictive tasks. We argue that this disagreement lies in how researchers process eye-tracking data as most remain agnostic to the human component and apply the data directly to DL models without proper preprocessing. We present EyeXNet, a multimodal DL architecture that combines images and radiologists’ fixation masks to predict abnormality locations in chest X-rays. We focus on fixation maps during reporting moments as radiologists are more likely to focus on regions with abnormalities and provide more targeted regions to the predictive models. Our analysis compares radiologist fixations in both silent and reporting moments, revealing that more targeted and focused fixations occur during reporting. Our results show that integrating the fixation masks in a multimodal DL architecture outperformed the baseline model in five out of eight experiments regarding average Recall and six out of eight regarding average Precision. Incorporating fixation masks representing radiologists’ classification patterns ina multimodal DL architecture benefits lesion detection in chest X-ray (CXR) images, particularly when there is a strong correlation between fixation masks and generated proposal regions. This highlights the potential of leveraging fixation masks to enhance multimodal DL architectures for CXR image analysis. This work represents a first step towards human-centered DL, moving away from traditional data-driven and human-agnostic approaches.  \nKeywords: multimodal deep learning; eye tracking; object detection; X-rays; fixation maps  \n1","cbCaivxaptYKlzjH","https://ap.wps.com/l/cbCaivxaptYKlzjH","pdf",9383204,1,17,"English","en",105,"# Introduction\n## Chest X-ray diagnosis and deep learning\n## Eye-tracking datasets and multimodal architectures\n## Conflicting evidence in the literature\n## Study motivation and objective","[{\"question\":\"Why do studies integrating eye gaze data with chest X-rays reach conflicting conclusions?\",\"answer\":\"The disagreement stems from how eye-tracking data are processed: many works apply the data directly to models while remaining agnostic to the human component and using insufficient preprocessing.\"},{\"question\":\"What is EyeXNet and how does it use eye-tracking information?\",\"answer\":\"EyeXNet is a multimodal deep learning architecture that combines chest X-ray images with radiologists’ fixation masks to predict abnormality locations, emphasizing fixation maps during reporting moments.\"},{\"question\":\"How do fixation masks affect model performance compared with a baseline?\",\"answer\":\"In eight experiments, incorporating fixation masks outperformed the baseline in five experiments for average Recall and in six for average Precision, particularly when fixation masks correlate with proposal regions.\"}]","EyeXNet - Enhancing Abnormality Detection and Diagnosis via Eye-Tracking and X-ray Fusion - Abstract | PDF",1785894764,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"eyexnet-enhancing-abnormality-detection-and-diagnosis-via-eye-tracking-and-x-ray-fusion-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/eyexnet-enhancing-abnormality-detection-and-diagnosis-via-eye-tracking-and-x-ray-fusion-abstract/124807/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do studies integrating eye gaze data with chest X-rays reach conflicting conclusions?","Question",{"text":75,"@type":76},"The disagreement stems from how eye-tracking data are processed: many works apply the data directly to models while remaining agnostic to the human component and using insufficient preprocessing.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is EyeXNet and how does it use eye-tracking information?",{"text":80,"@type":76},"EyeXNet is a multimodal deep learning architecture that combines chest X-ray images with radiologists’ fixation masks to predict abnormality locations, emphasizing fixation maps during reporting moments.",{"name":82,"@type":73,"acceptedAnswer":83},"How do fixation masks affect model performance compared with a baseline?",{"text":84,"@type":76},"In eight experiments, incorporating fixation masks outperformed the baseline in five experiments for average Recall and in six for average Precision, particularly when fixation masks correlate with proposal regions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]