[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128473-en":3,"doc-seo-128473-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128473,13056712833777,"Logic","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Optimisation for Efficient Deep Learning - Doctor of Philosophy thesis","Over the past decade, deep neural networks have dramatically increased performance across supervised learning, reshaping state-of-the-art results on major machine vision and natural language processing benchmarks. This progress has also increased deployment costs due to larger computation and higher energy usage. The thesis presents training- and deployment-focused optimisation methods that reduce cost while maintaining high accuracy. It introduces two tuning-light optimisers with fixed maximal step size, including a novel bundle method for interpolation and a Polyak-like step approach for non-interpolating settings, plus practical optimisation for fully binary networks.","Optimisation for Efficient Deep Learning  \nAlasdair James Paren  \nSt Anne’s College  \nA thesis presented for the degree of Doctor of Philosophy  \nDepartment of Engineering Science University of Oxford Trinity 2022  \nAcknowledgements  \nI would first like to thank Toshiba Research Laboratory Cambridge for partially funding my doctorate, which would not have been possible without their financial support.  \nI would next like to acknowledge my dedicated supervisors Pawan and Rudra for their patience throughout my time at the University of Oxford. Rudra provided me with great advice and encouragement, on both technical and non-technical matters. I’d like to thank Pawan for his engaging teaching, very thorough proofreading and attention to detail. I am especially grateful for Pawan’s decision to keep supervising me even after leaving his role at Oxford.  \nI would also like to thank my lab mates Rudy, Alban, Prateek, Jodie, Florian and Alessandro; it was a real pleasure to work alongside you and discuss all sorts of interesting topics and ideas with you. I’m sad COVID stole further opportunities to interact with all of you in person.  \nI would like to let my parents Julian and Mary know how much I appreciate their support. This including their investment in my education from an early age and their continual perspective and kind words of encouragement.  \nFinally, I would like to extend a special thank you to my partner, Maria. I’m not going to say “I could not have done it without you”, but it certainly would have taken me even longer and I would not have enjoyed it as much. Thank you for the walks, proof reading and cooking around deadlines. Thank you for your continued love and support, it has been a privilege to spend this time with you!  \nAbstract  \nOver the past 10 years there has been a huge advance in the performance power of deep neural networks on many supervised learning tasks. Over this period these models have redefined the state of the art numerous times on many classic machine vision and natural language processing benchmarks. Deep neural networks have also found their way into many real-world applications including chat bots, art generation, voice activated virtual assistants, surveillance, and medical diagnosis systems. Much of the improved performance of these models can be attributed to an increase in scale, which in turn has raised computation and energy costs.  \nIn this thesis we detail approaches of how to reduce the cost of deploying deep neural networks in various settings. We first focus on training efficiency, and to that end we present two optimisation techniques that produce high accuracy models without extensive tuning. These optimisers only have a single fixed maximal step size hyperparameter to cross-validate and we demonstrate that they outperform other comparable methods in a wide range of settings. These approaches do not require the onerous process of finding a good learning rate schedule, which often requires training many versions of the same network, hence they reduce the computation needed. The first of these optimisers is a novel bundle method designed for the interpolation setting. The second demonstrates the effectiveness of a Polyak-like step size in combination with an online estimate of the optimal loss value in thenon-interpolating setting.  \nNext, we turn our attention to training efficient binary networks with both binary parameters and activations. With the right implementation, fully binary networks are highly efficient at inference time, as they can replace the majority of operations with cheaper bit-wise alternatives. This makes them well suited for lightweight or embedded applications. Due to the discrete nature of these models conventional training approaches are not viable. We present a simple and effective alternative to the existing optimisation techniques for these models.  \nContents  \nContents iii  \nList of Figures 1  \nList of Tables 7  \n1 Introduction 10  \n1.1 Motivation ......","cbCaibSZepRaScp0","https://ap.wps.com/l/cbCaibSZepRaScp0","pdf",1365379,1,192,"English","en",105,"# Acknowledgements\n# Abstract\n# List of Figures\n# List of Tables\n# 1 Introduction\n## Motivation\n## Deep Neural Networks\n## Learning as Optimisation\n## Optimisation Algorithms\n## Thesis Outline and Contributions\n## Publications\n# 2 Related Work\n## Introduction\n## Stochastic Gradient Descent\n## Momentum\n## Line Search Methods\n## Adaptive Gradient Methods\n## Methods for Interpolation\n# 3 Preliminaries\n## Learning Task\n## Loss Function\n## Regularisation\n## Problem Formulation\n## Interpolation\n## First Order Methods\n## Adaptive Moment Estimation (Adam)\n## Adaptive Learning-rates for Interpolation with Gradients (ALI-G)\n# 4 Bundle Optimisation for Robust and Accurate Training (BORAT)\n## Introduction\n## Bundle Methods\n## The BORAT Algorithm\n## Advantages of Bundles with More Than Two Pieces\n## Primal Problem\n## Dual Problem\n## Selecting Additional Linear Approximations for the Bundle\n## Efficient Dual Algorithm to Compute N ≥ 2 Linear Pieces\n## Computational Considerations\n## Summary of the Algorithm","[{\"question\":\"What problem does the thesis address in efficient deep learning?\",\"answer\":\"It addresses the high computation and energy costs that come with scaling deep neural networks for high-performing supervised learning. It focuses on reducing deployment and training cost without sacrificing accuracy.\"},{\"question\":\"What are the two main optimisation approaches introduced for training efficiency?\",\"answer\":\"The thesis presents two optimisation techniques that aim for high accuracy with minimal tuning. One is a novel bundle method for the interpolation setting, and the other combines a Polyak-like step size with an online estimate of the optimal loss value for a non-interpolating setting.\"},{\"question\":\"Why are conventional training methods difficult for fully binary networks, and what solution is proposed?\",\"answer\":\"Fully binary networks are discrete in parameters and activations, making conventional gradient-based training approaches not viable. The thesis proposes a simple and effective optimisation alternative tailored to these models.\"}]","Optimisation for Efficient Deep Learning - Doctor of Philosophy thesis | PDF",1786001266,484,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"optimisation-for-efficient-deep-learning-doctor-of-philosophy-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/optimisation-for-efficient-deep-learning-doctor-of-philosophy-thesis/128473/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address in efficient deep learning?","Question",{"text":76,"@type":77},"It addresses the high computation and energy costs that come with scaling deep neural networks for high-performing supervised learning. It focuses on reducing deployment and training cost without sacrificing accuracy.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are the two main optimisation approaches introduced for training efficiency?",{"text":81,"@type":77},"The thesis presents two optimisation techniques that aim for high accuracy with minimal tuning. One is a novel bundle method for the interpolation setting, and the other combines a Polyak-like step size with an online estimate of the optimal loss value for a non-interpolating setting.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are conventional training methods difficult for fully binary networks, and what solution is proposed?",{"text":85,"@type":77},"Fully binary networks are discrete in parameters and activations, making conventional gradient-based training approaches not viable. The thesis proposes a simple and effective optimisation alternative tailored to these models.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]