[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126110-en":3,"doc-seo-126110-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126110,5909887254083,"Miles","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Phrase-Based Image Captioning","Generating novel textual descriptions of images bridges computer vision and natural language processing. This paper introduces a simple phrase-focused model that generates descriptive sentences from a sample image, emphasizing description syntax. It trains a purely bilinear model to learn a metric between image representations from a pre-trained CNN and phrase sets used in captions. From caption syntax statistics, a trigram constrained language model decodes relevant phrases and beam-search sentences, achieving competitive results on Flickr30k and Microsoft COCO.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk)  brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nID IAP RESEARCH REPORT  \nPHRASE-BASED IMAGE CAPTIONING  \nRémi Lebret a Pedro H. O. Pinheiro Ronan Collobert  \nIdiap-RR-08-2015  \nMAY 2015  \na Idiap  \nCentre du Parc , Rue Marconi 19, P.O. Box 592, CH-1920 Martigny  \nT +41 27 721 77 11 F +41 27 721 77 12 [info@idiap.ch](info@idiap.ch) [www.idiap.ch](www.idiap.ch)  \nPhrase-based Image Captioning  \nRmi Lebret􀀃 & Pedro O. Pinheiro􀀃 REMI @LEBRET. CH , PEDRO @ OPINHEIRO . COM  \nIdiap Research Institute, Martigny, Switzerland  \n´Ecole Polytechnique Fdrale de Lausanne (EPFL), Lausanne, Switzerland  \nRonan Collobert RONAN @COLLOBERT. COM  \nFacebook AI Research, Menlo Park, CA, USA  \nAbstract  \nGenerating a novel textual description of an image is an interesting problem that connects computer vision and natural language processing. In this paper, we present a simple model that is able to generate descriptive sentences given a sample image. This model has a strong focus on the syntax of the descriptions. We train a purely bilinear model that learns a metric between an image representation (generated from a previously trained Convolutional Neural Network) and phrases that are used to described them. The system is then able to infer phrases from a given image sample. Based on caption syntax statistics, we propose a simple language model that can produce relevant descriptions for a given test image using the phrases inferred. Our approach, which is considerably simpler than state-of-the-art models, achieves comparable results in two popular datasets for the task: Flickr30k and the recently proposed Microsoft COCO.  \n1. Introduction  \nBeing able to automatically generate a description from animage is a fundamental problem in artiﬁcial intelligence, connecting computer vision and natural language processing. The problem is particularly challenging because it requires to correctly recognize different objects in images and how they interact. Another challenge is that an image description generator needs to express these interactions ina natural language (e.g. English) . Therefore, a language model is implicitly required in addition to visual understanding.  \n􀀃 These two authors contributed equally to this work.  \nProceedings of the 32 nd International Conference on Machine Learning, Lille, France, 2015 . JMLR: W&CP volume 37 . Copyright 2015 by the author(s) .  \nRecently, this problem has been studied by many different authors. Most of the attempts are based on recurrent neural networks to generate sentences. These models leverage the power of neural networks to transform image and sentence representations into a common space (Mao et al., 2015 ; Karpathy & Fei-Fei, 2015 ; Vinyals et al., 2014 ; Donahue et al., 2014) .  \nIn this paper, we propose a different approach to the problem that does not rely on complex recurrent neural networks. An exploratory analysis of two large datasets of image descriptions reveals that their syntax is quite simple. The ground-truth descriptions can be represented asa collection of noun, verb and prepositional phrases. The different entities in a given image are described by the noun phrases, while the interactions or events between these entities are encoded by both the verb and the prepositional phrases. We thus train a model that predicts theset of phrases present in the sentences used to describe the images. By leveraging previous works on word vector representations, each phrase can be represented by the mean of the representations of the words that compose the phrase. Vector representations for images can also be easily obtained from some pre-trained convolutional neural networks. The model then learns a common embedding between phrase and image representations (see Figure 3) .  \nGiven a test image, a bilinear model is trained to predict a set of top-ranked phrase","cbCaivNj4zXYY19g","https://ap.wps.com/l/cbCaivNj4zXYY19g","pdf",1223216,1,12,"English","en",105,"# Introduction\n# Related Works\n# Phrase-Based Model\n# Sentence Generation\n# Experimental Setup and Results\n# Conclusion","[{\"question\":\"What problem does phrase-based image captioning address?\",\"answer\":\"It generates descriptive sentences for a given image by connecting visual recognition with language generation. The approach targets producing syntactically correct descriptions.\"},{\"question\":\"How does the proposed model represent phrases and images?\",\"answer\":\"Phrases are represented using the mean of word vector representations forming each phrase, while image representations come from a previously trained convolutional neural network. A bilinear model learns an embedding/metric between phrase and image representations.\"},{\"question\":\"How are sentences generated from the predicted phrases?\",\"answer\":\"A trigram constrained language model uses syntax statistics from the training captions to constrain decoding. Beam search infers candidate sentences from top-ranked phrase subsets, followed by a re-ranking step based on closeness to the image in the learned metric.\"}]","Phrase-Based Image Captioning | PDF",1785903225,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"phrase-based-image-captioning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/phrase-based-image-captioning/126110/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does phrase-based image captioning address?","Question",{"text":75,"@type":76},"It generates descriptive sentences for a given image by connecting visual recognition with language generation. The approach targets producing syntactically correct descriptions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed model represent phrases and images?",{"text":80,"@type":76},"Phrases are represented using the mean of word vector representations forming each phrase, while image representations come from a previously trained convolutional neural network. A bilinear model learns an embedding/metric between phrase and image representations.",{"name":82,"@type":73,"acceptedAnswer":83},"How are sentences generated from the predicted phrases?",{"text":84,"@type":76},"A trigram constrained language model uses syntax statistics from the training captions to constrain decoding. Beam search infers candidate sentences from top-ranked phrase subsets, followed by a re-ranking step based on closeness to the image in the learned metric.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]