[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122006-en":3,"doc-seo-122006-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122006,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Exploring Loss Functions in Machine Learning","The loss function is central to machine learning because it governs training, evaluation, and optimization, shaping how effectively and efficiently models solve specific tasks. The work introduces three new loss functions and demonstrates their use across classification and representation learning. It presents a tree loss as a cross-entropy replacement for multi-class problems with semantically similar categories. It also develops DocSplit, a contrastive pretraining method for large document embeddings, and proposes weighted contrastive loss to leverage inter-label relationships, improving fine-grained class distinction and fine-tuning performance on emotion and sentiment tasks.","Claremont Colleges  \nScholarship @ Claremont  \n\n| CGU Theses & Dissertations | CGU Student Scholarship |\n| --- | --- |\n| Summer 2024\u003Cbr>Exploring Loss Functions in Machine Learning\u003Cbr>Yujie Wang\u003Cbr>Claremont Graduate University\u003Cbr>Follow this and additional works at: [https://scholarship.claremont.edu/cgu_etd](https://scholarship.claremont.edu/cgu_etd)\u003Cbr> Part of the Mathematics Commons |  |\n\nRecommended Citation  \nWang, Yujie. (2024) . Exploring Loss Functions in Machine Learning. CGU Theses & Dissertations, 831. [https://scholarship.claremont.edu/cgu_etd/831](https://scholarship.claremont.edu/cgu_etd/831) .  \nThis Open Access Dissertation is brought to you for free and open access by the CGU Student Scholarship at Scholarship @ Claremont. It has been accepted for inclusion in CGU Theses & Dissertations by an authorized administrator of Scholarship @ Claremont. For more information, please contact [scholarship@claremont.edu](scholarship@claremont.edu).  \nExploring Loss Functions in Machine Learning  \nBy  \nYujie Wang  \nClaremont Graduate University  \n2024  \nCopyright Yujie Wang, 2024 All rights reserved.  \nApproval of the Dissertation Committee  \nThis dissertation has been duly read, reviewed, and critiqued by the Committee listed below, which hereby approves the manuscript of Yujie Wang as fulﬁlling the scope and quality requirements for meriting the degree of Doctor of Philosophy in Mathematics.  \nMike Izbicki, chair  \nClaremont McKenna College  \nAssistant Professor of Computer Science  \nJohn Angus  \nClaremont Graduate Unicersity  \nProfessor of Mathematics  \nQidi Peng  \nClaremont Graduate Unicersity  \nResearch Associate Professor of Mathematics  \nYu Bai  \nClaremont Graduate Unicersity  \nResearch Associate Professor of Mathematics  \nAbstract  \nExploring Loss Functions in Machine Learning  \nby  \nYujie Wang  \nClaremont Graduate University: 2024  \nThe loss function plays a critical role in machine learning. It is fundamental in training, evaluating, and optimizing machine learning models, directly impacting their effectiveness andefﬁciency in solving speciﬁc tasks. We explore three new loss functions and their applications. Softmax Cross-Entropy Loss, stands as a prevalent choice in neural network classiﬁcation tasks. It treats all misclassiﬁcations uniformly. However, multi-class classiﬁcation problems often have many semantically similar classes. We should expect that these semantically similar classes will have similar parameter vectors. We introduce a weighted loss function, the tree loss as a drop-in replacement for the cross entropy loss. The tree loss re-parameterizes the parameter matrix in order to guarantee that semantically similar classes will have similar parameter vectors. Using simple properties of stochastic gradient descent, we show that the the tree loss’s generalization error is asymptotically better than the cross entropy loss’s. We then validate these theoretical results on synthetic data, image data (CIFAR100, ImageNet), and text data (Twitter) .  \nWe also investigate the application of contrastive loss in large document embeddings. Existing model pretraining methods primarily focus on local information. For instance, in the widely used token masking strategy, words closer to the masked token are given more  \nimportance for prediction than words further away. While this approach results in pretrained models that generate high-quality sentence embeddings, it leads to low-quality embeddings for larger documents. We propose a new pretraining method called DocSplit, which compels models to consider the entire global context of a large document. Our method employs a contrastive loss where the positive examples are randomly sampled sections of the input document, and  \nthe negative examples are randomly sampled sections from unrelated documents. Similar to previous pretraining methods, DocSplit is fully unsupervised, straightforward to implement, and can be used to pretrain any model architecture. Our experiment","cbCaihlblOosuTZq","https://ap.wps.com/l/cbCaihlblOosuTZq","pdf",2127592,1,75,"English","en",105,"# Chapter 1 Introduction\n## 1.1 Weighted Loss Function\n## 1.2 Contrastive Loss Function\n## 1.3 Weighted Contrastive Loss Function\n# Chapter 2 The Tree Loss: Improving Generalization with Many Classes\n## 2.1 Introduction\n## 2.2 Problem Setting\n## 2.3 The Tree Loss\n## 2.3.1 The U-Tree Loss\n## 2.3.2 The V-Tree Loss\n## 2.3.3 Intuition","[{\"question\":\"Why are loss functions critical in machine learning?\",\"answer\":\"Loss functions determine how models are trained, evaluated, and optimized. They directly influence a model’s effectiveness and efficiency in completing specific tasks.\"},{\"question\":\"What problem does tree loss address compared with softmax cross-entropy?\",\"answer\":\"Softmax cross-entropy treats all misclassifications uniformly, but multi-class datasets often contain semantically similar classes. Tree loss re-parameterizes class relationships so semantically similar classes obtain similar parameter vectors.\"},{\"question\":\"How does DocSplit improve embeddings for large documents?\",\"answer\":\"DocSplit uses a contrastive loss during unsupervised pretraining. It samples positive examples as sections from the same document and negatives from unrelated documents, forcing the model to integrate global context.\"}]","Exploring Loss Functions in Machine Learning | PDF",1785808257,189,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"exploring-loss-functions-in-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/exploring-loss-functions-in-machine-learning/122006/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are loss functions critical in machine learning?","Question",{"text":75,"@type":76},"Loss functions determine how models are trained, evaluated, and optimized. They directly influence a model’s effectiveness and efficiency in completing specific tasks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does tree loss address compared with softmax cross-entropy?",{"text":80,"@type":76},"Softmax cross-entropy treats all misclassifications uniformly, but multi-class datasets often contain semantically similar classes. Tree loss re-parameterizes class relationships so semantically similar classes obtain similar parameter vectors.",{"name":82,"@type":73,"acceptedAnswer":83},"How does DocSplit improve embeddings for large documents?",{"text":84,"@type":76},"DocSplit uses a contrastive loss during unsupervised pretraining. It samples positive examples as sections from the same document and negatives from unrelated documents, forcing the model to integrate global context.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]