[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120243-en":3,"doc-seo-120243-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120243,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Understanding Encoder-Decoder Structures in Machine Learning Using Information Measures - Abstract and Key Results","The document presents new information-theoretic results for modeling and understanding encoder-decoder design in machine learning. It introduces information sufficiency (IS) and mutual information loss (MIL) to characterize predictive structures and give a functional expression for probabilistic models consistent with an IS encoder-decoder latent structure. The work justifies forward stages in many modern architectures, analyzes performance loss via cross-entropy risk, and shows MIL quantifies expressiveness limitations from biased designs. It also establishes necessary and sufficiency conditions for universal cross-entropy learning using encoder-decoder designs.","arXiv :2405 .20452v1 [ cs .LG] 30 May 2024  \nUnderstanding Encoder-Decoder Structures in Machine Learning Using Information Measures  \nJorge F. Silva, Victor Faraggi, Camilo Ramirez, Alvaro Ega˜na y Eduardo Pavez  \nAbstract  \nWe present new results to model and understand the role of encoder-decoder design in machine learning (ML) from an information-theoretic angle. We use two main information concepts, information sufficiency (IS) and mutual information loss (MIL), to represent predictive structures in machine learning. Our first main result provides a functional expression that characterizes the class of probabilistic models consistent with an IS encoder-decoder latent predictive structure. This result formally justifies the encoder-decoder forward stages many modern ML architectures adopt to learn latent (compressed) representations for classification. To illustrate IS as a realistic and relevant model assumption, we revisit some known ML concepts and present some interesting new examples: invariant, robust, sparse, and digital models. Furthermore, our IS characterization allows us to tackle the fundamental question of how much performance (predictive expressiveness) could be lost, using the cross entropy risk, when a given encoderdecoder architecture is adopted in a learning setting. Here, our second main result shows that a mutual information loss quantifies the lack of expressiveness attributed to the choice of a (biased) encoder-decoder ML design. Finally, we address the problem of universal cross-entropy learning with an encoder-decoder design where necessary and sufficiency conditions are established to meet this requirement. In all these results, Shannon’s information measures offer new interpretations and explanations for representation learning.  \nIndex Terms  \nRepresentation learning, learning and coding, encoder-decoder design, explainability, encoder expressiveness, Shannon information measures, information sufficiency, sparse models, digital models, invariant models, information bottleneck.  \nI. INTRODUCTION  \nIn many machine learning (ML) tasks, the observation (or input) X lives in a continuous (multivariate) high dimensional space while the class (target) variable Y is discrete. Given this discrepancy at the space level, it is common to assume that there are many latent factors (random innovation components) that produce X but do not affect Y and, consequently, this redundant information is not needed for learning a good classifier [1] . Therefore, representation learning (RL) addresses the compression task of finding a lossy transformation of an observation X (or encoder) that is highly informative and ideally sufficient for predicting Y. A large body of work addresses the design of lossy representations (compressors) from data. Many of these methods rely on the use of information-theoretic measures to quantify the predictive relationship between X and Y [2]–[8] . This general idea of redundancy on X to explain Y translates intuitively into a notion of probabilistic structure [9]–[12] that is one of the key justifications for the adoption of encoder-decoder strategies in ML.  \nThe formal characterization of probabilistic structures and the study of the repercussions of these model assumptions in the design of algorithms and architectures is an essential area of theoretical research in ML [9], [10], [12]–[15] . This theoretical understanding has been used to explain some design choices of neural network architectures [9], [16] and has been adopted to design data compressors for prediction [13] .  \nOn formalizing probabilistic structures in learning and connecting it with the idea of sufficient representations for inference (classification), we highlight the paper by Bloem-Reddy and Teh [9] . This seminal work studies probabilistic models (where a model is a joint distribution µX,Y between X and Y) that are invariant to the action of a compact group of measurable transformations G. Classification tasks invar","cbCaieq1tbOd7kd7","https://ap.wps.com/l/cbCaieq1tbOd7kd7","pdf",1459568,1,32,"English","en",105,"# Introduction\n## Latent factors and representation learning\n## Probabilistic structures and invariance\n## Connection to sufficient encoders\n# Contributions\n## Extending representation learning theory","[{\"question\":\"What information-theoretic concepts does the paper use to study encoder-decoder structures?\",\"answer\":\"It uses information sufficiency (IS) and mutual information loss (MIL) to represent and analyze predictive structures in machine learning.\"},{\"question\":\"How does the paper justify encoder-decoder forward stages in modern ML architectures?\",\"answer\":\"It provides a characterization of probabilistic models consistent with an IS encoder-decoder latent predictive structure, formally supporting the use of encoder-decoder stages to learn latent compressed representations for classification.\"},{\"question\":\"How is performance or predictive expressiveness loss quantified when using a given encoder-decoder architecture?\",\"answer\":\"The paper relates expressiveness loss to cross-entropy risk, and its second main result shows that mutual information loss quantifies the lack of expressiveness caused by biased encoder-decoder designs.\"}]","Understanding Encoder-Decoder Structures in Machine Learning Using Information Measures - Abstract and Key Results | PDF",1785728952,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"understanding-encoder-decoder-structures-in-machine-learning-using-information-measures-abstract-and-key-results","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/understanding-encoder-decoder-structures-in-machine-learning-using-information-measures-abstract-and-key-results/120243/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What information-theoretic concepts does the paper use to study encoder-decoder structures?","Question",{"text":75,"@type":76},"It uses information sufficiency (IS) and mutual information loss (MIL) to represent and analyze predictive structures in machine learning.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper justify encoder-decoder forward stages in modern ML architectures?",{"text":80,"@type":76},"It provides a characterization of probabilistic models consistent with an IS encoder-decoder latent predictive structure, formally supporting the use of encoder-decoder stages to learn latent compressed representations for classification.",{"name":82,"@type":73,"acceptedAnswer":83},"How is performance or predictive expressiveness loss quantified when using a given encoder-decoder architecture?",{"text":84,"@type":76},"The paper relates expressiveness loss to cross-entropy risk, and its second main result shows that mutual information loss quantifies the lack of expressiveness caused by biased encoder-decoder designs.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]