[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119605-en":3,"doc-seo-119605-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119605,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Valence and Arousal prediction from food images using Deep Learning and classical Machine Learning models - Research overview","A computational framework predicts emotional dimensions—valence and arousal—from visual representations of food imagery. The work analyzes how color, brightness, texture, and spatial composition contribute to affective perception, and models these cues quantitatively with machine learning. Using 1,211 annotated food images with continuous valence–arousal labels, high-level embeddings are extracted with a pre-trained Vision Transformer (ViT). Three regression models (RF, SVR, MLP) are trained under matched preprocessing and evaluated with MSE, MAE, and R2, where MLP provides the most stable convergence and best accuracy.","DIPARTIMENTO DI INGEGNERIA DELL’INFORMAZIONE  \nCORSO DI LAUREA MAGISTRALE IN ICT FOR INTERNET AND MULTIMEDIA  \nValence and Arousal prediction from food images using Deep Learning and classical Machine Learning models  \nSupervisor: Prof. Antonio Roda  \nCo-supervisor: Dr. Matteo Spanio  \nLaureando: Fatemeh Talebi  \nNovember 17, 2025  \nAbstract  \nIn this work, we present a computational framework for predicting emotional dimensions—valence and arousal—from visual representations of food imagery. The study aims to investigate how visual cues such as color, brightness, texture, and spatial composition influence the affective perception of food, and how these cues can be quantitatively modeled using machine learning techniques. Leveraging a dataset of 1,211 annotated food images with continuous valence–arousal labels, we extracted high-level visual embeddings using a pre-trained Vision Transformer (ViT)[1] model, which effectively captures semantic and aesthetic characteristics of the images.  \nThree regression models—Random Forest (RF), Support Vector Regression (SVR), and a Multi-Layer Perceptron (MLP)—were implemented to evaluate the relationship between extracted features and emotional responses. All models were trained under identical preprocessing and standardization procedures to ensure fairness and reproducibility. Quantitative evaluation was performed using Mean Squared Error (MSE), Mean Absolute Error (MAE), and Coefficient of Determination (R2 ) metrics. The experimental results revealed that the MLP model achieved the most stable convergence and the highest predictive accuracy, outperforming classical regressors. Extending the training duration from  \n100 to 500 epochs further improved performance consistency, demonstrating that longer optimization can enhance emotional representation learning without causing overfitting.  \nNevertheless, several challenges were observed, including limited datasetsize, uneven emotional intensity distribution, and model sensitivity to ambiguous or low-contrast food images. These factors occasionally led to larger prediction errors, especially for samples with moderate emotional content. Despite these limitations, the findings underscore the potential of deep learning models—particularly neural architectures combined with transformer-based feature extraction—to effectively capture the nuanced relationship between food aesthetics and affective perception. This research establishes a reproducible baseline for visual emotion prediction and opens pathways for future work on larger datasets, multimodal fusion, and interpretable emotion modeling in affective computing.  \nKeywords: Affective Computing, Emotion Prediction, Food Images, Valence–Arousal Model, Vision Transformer (ViT), Random Forest, Support Vector Regression, Multi-Layer Perceptron  \nContents  \nAbstract 2  \n1 Introduction 5  \n1.1 Background and Context ...................... 5  \n1.2 Visual Emotion Analysis and the Valence–Arousal Space ..... 5  \n1.3 Food Imagery and Emotional Perception .............. 5  \n1.4 Machine Learning for Emotion Prediction ............. 6  \n1.5 Dataset and Annotation Framework ................ 6  \n1.6 Motivation and Significance ..................... 6  \n1.7 Motivation of Study ......................... 7  \n1.8 Problem Statement .......................... 7  \n1.9 Research Objectives ......................... 7  \n2 Related Work 8  \n2.1 Emotion Recognition in Images ................... 8  \n2.1.1 From Facial Expression Analysis to General Visual Emotion 8  \n2.1.2 Transition to Affective Image Understanding ....... 8  \n2.1.3 Dimensional Emotion Representation ........... 9  \n2.1.4 Feature Evolution: From Handcrafted to Deep Representations ............................. 9  \n2.1.5 Challenges in Visual Emotion Prediction .......... 9  \n2.2 Visual Emotion Datasets ....................... 10  \n2.2.1 Early Emotion Datasets: From Psychology to Visual Media 10  \n2.2.2 Affective Datasets for Artistic and Natural Images ","cbCaivwr1xAIm5dU","https://ap.wps.com/l/cbCaivwr1xAIm5dU","pdf",3756088,1,68,"English","en",105,"# Introduction\n## Background and Context\n## Visual Emotion Analysis and the Valence–Arousal Space\n## Food Imagery and Emotional Perception\n## Machine Learning for Emotion Prediction\n## Dataset and Annotation Framework\n## Motivation and Significance\n# Related Work\n## Emotion Recognition in Images\n## Visual Emotion Datasets\n## Emotion Representation Models and Theoretical Foundations\n## Deep Learning Approaches for Emotion Recognition\n## Applications of Visual Emotion Recognition\n# Methodology\n## Dataset Description and Preprocessing\n## Model Architecture and Training Configuration\n## Evaluation Metrics and Experimental Setup\n# Experiments\n## Tools and Frameworks","[{\"question\":\"What emotional targets does the framework predict from food images?\",\"answer\":\"It predicts valence and arousal, two continuous emotional dimensions, from the visual content of food imagery.\"},{\"question\":\"How are image features obtained for emotion prediction?\",\"answer\":\"A pre-trained Vision Transformer (ViT) extracts high-level visual embeddings that capture semantic and aesthetic characteristics.\"},{\"question\":\"Which regression model performs best and how is it evaluated?\",\"answer\":\"The Multi-Layer Perceptron (MLP) achieves the most stable convergence and the highest predictive accuracy, evaluated using MSE, MAE, and R2.\"}]","Valence and Arousal prediction from food images using Deep Learning and classical Machine Learning models - Research overview | PDF",1785725259,171,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"valence-and-arousal-prediction-from-food-images-using-deep-learning-and-classical-machine-learning-models-research-overview","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/valence-and-arousal-prediction-from-food-images-using-deep-learning-and-classical-machine-learning-models-research-overview/119605/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What emotional targets does the framework predict from food images?","Question",{"text":75,"@type":76},"It predicts valence and arousal, two continuous emotional dimensions, from the visual content of food imagery.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are image features obtained for emotion prediction?",{"text":80,"@type":76},"A pre-trained Vision Transformer (ViT) extracts high-level visual embeddings that capture semantic and aesthetic characteristics.",{"name":82,"@type":73,"acceptedAnswer":83},"Which regression model performs best and how is it evaluated?",{"text":84,"@type":76},"The Multi-Layer Perceptron (MLP) achieves the most stable convergence and the highest predictive accuracy, evaluated using MSE, MAE, and R2.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]