[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85143-en":3,"doc-seo-85143-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85143,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Calibrated Hybrid CNN-Transformer for Retinal OCT Classification","Deep models for retinal optical coherence tomography (OCT) classification deliver high accuracy but seldom verify whether reported confidence is trustworthy, risking delays in sight-saving care. This work combines a hybrid convolutional–Transformer encoder with an XGBoost classification head, plus a clinical safety layer consisting of confidence calibration, out-of-distribution rejection, and per-prediction uncertainty flagging. On four-class OCT (84,495 scans) it achieves 95.4% accuracy and reduces expected calibration error twelve-fold (ECE = 0.0024), aligning confidence with true accuracy. The study jointly validates all three safety mechanisms with public weights and reproducible multi-seed evaluation.","Calibrated Hybrid CNN–Transformer for Retinal OCT Classification  \nAnimesh Kumar  \nSchool of Computing, Newcastle University, UK  \n[A.Kumar12@newcastle.ac.uk](A.Kumar12@newcastle.ac.uk) ORCID: 0009-0003-0608-7004  \nMarch 2026  \narXiv :2607 .09809v1 [ ee ss .IV] 9 Jul 2026  \nAbstract. Deep models for retinal optical coherence tomography (OCT) classification report high accuracy but rarely report whether their confidence can be trusted—a gap that matters when a wrong-but-confident reading delays sight-saving treatment. We pair a hybrid convolutional–Transformer encoder with a gradientboosting (XGBoost) classification head and a three-part clinical safety layer: confidence calibration, out-ofdistribution (OOD) rejection, and per-prediction uncertainty flagging. On four-class OCT (84,495 scans) the model reaches 95.4% accuracy while cutting calibration error twelve-fold (expected calibration error, ECE = 0 .0024), so the confidence it reports tracks its true accuracy. To our knowledge this is the first OCT classifier to validate all three safety mechanisms jointly, with public weights and reproducible multi-seed evaluation.  \nKeywords: retinal OCT, EfficientNetV2, vision transformer, XGBoost, OOD detection, calibration, GradCAM.  \n1 Introduction  \nOptical coherence tomography (OCT) is a non-invasive imaging technique that produces cross-sectional “B-scans”of the retina, and it is the standard modality for diagnosing four conditions that each demand a different clinical response: choroidal neovascularisation (CNV, the leaky abnormal vessels of “wet” age-related macular degeneration, requiring immediate anti-vascular endothelial growth factor [anti-VEGF] injection), diabetic macular oedema (DME), drusen (lipid deposits beneath the retinal pigment epithelium [RPE] that are an early biomarker of age-related macular degeneration, AMD), and healthy (Normal) tissue. Getting the class wrong—or getting it right at the wrong confidence level—carries a direct cost to a patient’s vision.  \nWhy this matters. Untreated CNV and DME are leading causes of preventable blindness, and the volume of OCT scans now far exceeds the number of specialists available to read them. Automated triage is therefore attractive, but a screening tool is only safe if a clinician can act on the confidence it reports. A model that is 95% accurate on average yet silently returns 95% confidence on a blurred, corrupted, or out-of-scope scan is actively dangerous in a triage setting. Accuracy alone does not make a model deployable; trustworthy confidence does. This paper targets that gap.  \nTwo problems have improved more slowly than raw accuracy since Kermany et al. [7] . First, convolutional neural networks (CNNs) process bounded receptive fields and cannot model the long-range dependencies that span  \nretinal layers. Second, softmax outputs are systematically overconfident [5] . We address both with a hybrid encoder for global context and a three-layer safety envelope for honest confidence.  \n2 Related Work  \nCNN-based OCT classification. Kermany et al. [7] showed that InceptionV3 with transfer learning performs strongly on four-class OCT data. Later work substituted VGG, ResNet, and DenseNet backbones [6, 9] . These improved accuracy in some settings but remained purely convolutional and included no safety components.  \nHybrid CNN–Transformer architectures. The vision transformer (ViT) [3] models global relationships between image patches through self-attention but needs large datasets. Hybrid architectures keep the CNN’s spatial inductive bias while adding global context—a practical trade-off for moderate-sized clinical collections such as Kermany.  \nGradient boosting as a classification head. XGBoost [2] fitted on deep feature vectors can sharpen class boundaries, especially for under-represented classes where a single softmax projection generalises poorly. We are not aware of prior OCT work combining a Transformer encoder with an XGBoost head.  \nClinical safety. Tem","cbCaicJVNQJzdHPN","https://ap.wps.com/l/cbCaicJVNQJzdHPN","pdf",1424288,3,1,4,"English","en",105,"# Introduction\n# Related Work\n# Methodology\n## Dataset\n## Architecture","[{\"question\":\"What main problem does the paper address in retinal OCT classification?\",\"answer\":\"It targets the gap between high average accuracy and unreliable confidence estimates, which can be dangerous during clinical triage when a model is overconfident on blurred, corrupted, or out-of-scope scans.\"},{\"question\":\"How does the proposed method improve trustworthiness of model outputs?\",\"answer\":\"It uses a confidence calibration step, out-of-distribution (OOD) rejection, and uncertainty flagging so that reported confidence better reflects true accuracy and inputs outside the training distribution are handled safely.\"},{\"question\":\"What results does the method achieve on the four-class OCT dataset?\",\"answer\":\"On four-class OCT (84,495 scans) the model reaches 95.4% accuracy and cuts calibration error twelve-fold, with expected calibration error (ECE) reported as 0.0024, indicating confidence tracks accuracy.\"}]",1784201356,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"calibrated-hybrid-cnn-transformer-for-retinal-oct-classification","",{"@graph":36,"@context":84},[37,52,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":22},"https://docshare.wps.com/document/calibrated-hybrid-cnn-transformer-for-retinal-oct-classification/85143/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":24,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":41,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What main problem does the paper address in retinal OCT classification?","Question",{"text":74,"@type":75},"It targets the gap between high average accuracy and unreliable confidence estimates, which can be dangerous during clinical triage when a model is overconfident on blurred, corrupted, or out-of-scope scans.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed method improve trustworthiness of model outputs?",{"text":79,"@type":75},"It uses a confidence calibration step, out-of-distribution (OOD) rejection, and uncertainty flagging so that reported confidence better reflects true accuracy and inputs outside the training distribution are handled safely.",{"name":81,"@type":72,"acceptedAnswer":82},"What results does the method achieve on the four-class OCT dataset?",{"text":83,"@type":75},"On four-class OCT (84,495 scans) the model reaches 95.4% accuracy and cuts calibration error twelve-fold, with expected calibration error (ECE) reported as 0.0024, indicating confidence tracks accuracy.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]