[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126479-en":3,"doc-seo-126479-105":31,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126479,962084925290,"Ophelia","https://ap-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","A localization/verification scheme for finding text in images and video frames based on contrast independent features and machine learning methods","Automatic character detection in video sequences is challenging because text varies in size and color while appearing over complex, cluttered backgrounds. This work presents a localization and verification scheme: candidate text regions are quickly localized with a low rejection rate to support character size normalization, then verified using machine-learning models trained on contrast-independent features. Multilayer perceptrons and support vector machines are compared using multiple feature sets. The approach enables fast text detection in images and videos with low computational cost compared with traditional methods.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nSignal Processing: Image Communication 19 (2004) 205–217  \nA localization/veriﬁcation scheme for ﬁnding text in images and video frames based on contrast independent features and  \nmachine learning methods  \nDatong Chena, *, Jean-Marc Odobeza , Jean-Philippe Thiranb  \naDalle Molle Institute for Perceptual Artiﬁcial Intelligence (IDIAP), Rue du Simplon 4, 1920 Martigny, Switzerland b Swiss Federal Institute of Technology Lausanne (EPFL), Signal Processing Institute (ITS), 1015 Lausanne, Switzerland  \nReceived 16 December 2002; received in revised form 6 May 2003; accepted 12 June 2003  \nAbstract  \nAutomatic character detection in video sequences is a complex task, due to the variety of sizes and colors as well as to the complexity of the background. In this paper we address this problem by proposing a localization/veriﬁcation scheme. Candidate text regions are ﬁrst localized by using a fast algorithm with a very low rejection rate, which enables the character size normalization. Contrast independent features are then proposed for training machine learning tools in order to verify the text regions. Two kinds of machine learning tools, multilayer perceptrons and support vector machines, are compared based on four different features in the veriﬁcation task. This scheme provides fast text detection in images and videos with a low computation cost, comparing with traditional methods.  \nr 2003 Elsevier B.V. All rights reserved.  \nKeywords: Image OCR; Video OCR; Content-based Indexing; Text detection; Support vector machine  \n1. Introduction  \nContent-based multimedia database indexing and retrieval tasks require automatic extraction of descriptive features that are relevant to the subject materials (images, video, etc.) . The typical low level features that are extracted in images or video include measures of color [25], texture [13], or shape [14] . Although these features can easily be extracted, they do not give a clear idea of the image content. Extracting more descriptive fea-  \n*Corresponding author.  \nE-mail addresses: chen@idiap.ch (D. Chen), odobez@ [idiap.ch](idiap.ch) (J.-M. Odobez), jp.thiran@epﬂ .ch (J.-P. Thiran) .  \ntures and higher level entities, for example text [3] or human faces [24], has attracted more and more research interest recently. Text embedded in images and video, especially captions, provide brief and important content information, such asthe name of players or speakers, the title, location and date of an event, etc. These text can be considered as a powerful feature (keyword) resource as are the information provided by speech recognizers for example. Besides, text-based search has been successfully applied in many applications while the robustness and computation cost of the feature matching algorithms based on other high level features are not efﬁcient enough to be applied to large databases.  \n0923-5965/$ -see front matter r 2003 Elsevier B.V. All rights reserved. doi:10.1016/S0923-5965(03)00075-4  \n206 D. Chen et al. / Signal Processing: Image Communication 19 (2004) 205–217  \nText detection and recognition in images and video frames, which aims at integrating advanced optical character recognition (OCR) and textbased searching technologies, is now recognized as a key component in the development of advanced image and video annotation and retrieval systems. Unfortunately, contrary to most of the OCR problems, text characters contained in images and videos can be any grayscale value (not always black characters on a white background), low resolution, variable size, and embedded in complex backgrounds. Experiments show that applying conventional OCR technology directly on the video frames leads to poor recognition rates. Therefore, efﬁcient detection and segmentation of text characters from background is ne","cbCaitsYoBYVnt3X","https://ap.wps.com/l/cbCaitsYoBYVnt3X","pdf",541272,2,1,13,"English","en",105,"# Introduction\n## Motivation: text as a high-level multimedia feature\n## Challenge: OCR on video frames under complex backgrounds\n## Prior work in text detection/localization\n## Contrast-independent feature verification approach","[{\"question\":\"Which machine learning models are compared for text region verification?\",\"answer\":\"The scheme compares multilayer perceptrons and support vector machines using four different features in the verification task.\"}]","A localization/verification scheme for finding text in images and video frames based on contrast independent features and machine learning methods | PDF",1785905281,33,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":29},"a-localizationverification-scheme-for-finding-text-in-images-and-video-frames-based-on-contrast-independent-features-and-machine-learning-methods","",{"@graph":37,"@context":78},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/a-localizationverification-scheme-for-finding-text-in-images-and-video-frames-based-on-contrast-independent-features-and-machine-learning-methods/126479/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-15","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"Which machine learning models are compared for text region verification?","Question",{"text":76,"@type":77},"The scheme compares multilayer perceptrons and support vector machines using four different features in the verification task.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,121,124,128],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":47,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":47,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":99,"slug":131},19,"General","general"]