[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118131-en":3,"doc-seo-118131-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118131,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","A Multimodal Machine Learning Approach to Generate News Articles from Geo-Tagged Images","This paper proposes an automated multimodal system that generates news articles from geo-tagged images. It first extracts location information from EXIF metadata, then uses convolutional neural networks with VGGNet16-based visual feature extraction to derive image representations and generate corresponding headlines. The headlines and visual features are further used to produce full articles using transformer-based text generation, including BLOOM. Experiments on real-world and online images validate outputs with ROUGE and BLEU scores, demonstrating effectiveness for journalism-oriented content creation.","A multimodal machine learning approach to generate news articles from geo-tagged images  \nAbhay Gotmare1, Gandharva Thite2, Laxmi Bewoor2  \n1Department of Information Technology, Vishwakarma Institute of Information Technology, Pune, India 2Department of Computer Engineering, Vishwakarma Institute of Information Technology, Pune, India  \nArticle history:  \nReceived Nov 30, 2023 Revised Mar 7, 2024 Accepted Mar 9, 2024  \nKeywords:  \nBLOOM  \nConvolutional neural network Journalism  \nLarge language model  \nLong short-term memory Multimodal machine learning VGGNet16  \nCorresponding Author:  \nClassical machine learning algorithms typically operate on unimodal data and hence it can analyze and make predictions based on data from a single source (modality) . Whereas multimodal machine learning algorithm, learns from information across multiple modalities, such as text, images, audio, and sensor data. The paper leverages the functionalities of multimodal machine learning (ML) application for generating text from images. The proposed work presents an innovative multimodal algorithm that automates the creation of news articles from geo-tagged images by leveraging cutting-edge developments in machine learning, image captioning, and advanced text generation technologies. Employing a multimodal approach that integrates machine learning and transformer algorithms, such as visual geometry group network16 (VGGNet16), convolutional neural network (CNN) and a long short-term memory (LSTM) based system, the algorithm initiates by extracting the location from exchangeable image file format (Exif) data from the image. The features are extracted from the image and corresponding news headline is generated. The headlines are used for generating a comprehensive article with contemporary large language model (LLM) . Further, the algorithm generates the news article big-science large openscience open-access multilingual language model (BLOOM) . The algorithm was tested on real time photographs as well as images from the internet. In both the cases the news articles generated were validated with ROUGE and BULE score. The proposed work is found to be successful attempt in journalism field.  \nThis is an open access article under the CC BY-SA license.  \nLaxmi Bewoor  \nDepartment of Computer Engineering, Vishwakarma Institute of Information Technology Pune, India  \nEmail: [laxmi.bewoor@viit.ac.in](laxmi.bewoor@viit.ac.in)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nThe manner that news is distributed and consumed has significantly changed with the advent of the digital age. Images have become a crucial component of this new environment because of their innate capacity to concisely express complex storylines. They frequently contain important information that supplements textual content and occasionally even replaces it. However, creating news stories manually from photos is a time-consuming task that requires a lot of labor and knowledge. Traditional machine learning (ML) algorithms have limitations with the uniform data format, but integrating more than one type of data is the need of the hour. Multimodal functionalities [1], [2] have recently garnered the attention of researchers to overcome this limitation. Multimodal applications have proven to be effective for hybridizing the models [3],[4] for text and image data types. The proposed work also represents an endeavor to develop a multimodal  \napplication and, in turn, step into the world of generative artificial intelligence (AI) [5], [6] . Therefore, the creation of automated systems capable of carrying out this activity could completely transform the journalism industry by increasing the productivity and scalability of news output. The development of artificial intelligence and machine learning technologies [7], [8] in recent years has opened significant opportunities for automating different parts of content generation. Particularly impressive improvements have been made in text generation models and ","cbCaibVYVOaQStm0","https://ap.wps.com/l/cbCaibVYVOaQStm0","pdf",540291,1,9,"English","en",105,"# Introduction\n## Problem Motivation\n## Multimodal Learning for Text-Image Generation\n## Transformers and Pretrained Language Models\n## Image Captioning Background\n## Gap and Research Goal","[{\"question\":\"How does the system start generating news content from an image?\",\"answer\":\"It extracts location information from the image’s EXIF data, then extracts visual features using a VGGNet16/CNN-based pipeline to support headline generation.\"},{\"question\":\"What models are used to produce headlines and full articles?\",\"answer\":\"A multimodal approach integrates CNN/VGGNet16 for visual feature extraction with transformer-based large language models, including BLOOM, to generate comprehensive news articles.\"},{\"question\":\"How are the generated news articles evaluated?\",\"answer\":\"The paper reports validation using ROUGE and BLEU scores on both real-time photographs and images collected from the internet.\"}]","A Multimodal Machine Learning Approach to Generate News Articles from Geo-Tagged Images | PDF",1785681780,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-multimodal-machine-learning-approach-to-generate-news-articles-from-geo-tagged-images","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-multimodal-machine-learning-approach-to-generate-news-articles-from-geo-tagged-images/118131/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the system start generating news content from an image?","Question",{"text":75,"@type":76},"It extracts location information from the image’s EXIF data, then extracts visual features using a VGGNet16/CNN-based pipeline to support headline generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What models are used to produce headlines and full articles?",{"text":80,"@type":76},"A multimodal approach integrates CNN/VGGNet16 for visual feature extraction with transformer-based large language models, including BLOOM, to generate comprehensive news articles.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the generated news articles evaluated?",{"text":84,"@type":76},"The paper reports validation using ROUGE and BLEU scores on both real-time photographs and images collected from the internet.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]