[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122857-en":3,"doc-seo-122857-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122857,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Privacy-Preserving Machine Learning on Apache Spark - Soteria Hybrid Trusted Execution Environments Scheme","Adoption of third-party machine learning cloud services depends on security assurances and the performance overhead introduced during model training and inference. The work studies security–performance trade-offs for distributed Apache Spark and its ML library. It leverages an insight that carefully selected non-sensitive operations, such as certain statistical calculations, can be revealed to improve privacy-preserving performance. The proposed Soteria system uses Trusted Execution Environments like Intel SGX and introduces a hybrid approach combining computation inside and outside enclaves, reducing runtime by up to 41% and providing a security proof against many ML attacks.","Received 12 October 2023, accepted 5 November 2023, date of publication 13 November 2023, date of current version 17 November 2023.  \nDigital Object Identifier 10.1109/ACCESS.2023.3332222  \nPrivacy-Preserving Machine Learning on Apache Spark  \nCLÁUDIA V. BRITO1,2, PEDRO G. FERREIRA1,3, BERNARDO L. PORTELA1,3, RUI C. OLIVEIRA 1,2, AND JOÃO T. PAULO1,2  \n1INESC TEC, 4200-465 Porto, Portugal  \n2Department of Informatics, University of Minho, 4710-057 Braga, Portugal  \n3Faculty of Sciences, University of Porto, 4099-002 Porto, Portugal Corresponding author: Cláudia V. Brito ([claudia.v.brito@inesctec.pt](claudia.v.brito@inesctec.pt))  \nThis work was supported by FCT-Portuguese Foundation for Science and Technology through the Ph.D. grant DFA/BD/146528/2018 and realized within the scope of the project LA/P/0063/2020 .  \nABSTRACT The adoption of third-party machine learning (ML) cloud services is highly dependent on the security guarantees and the performance penalty they incur on workloads for model training and inference. This paper explores security/performance trade-offs for the distributed Apache Spark framework and its ML library. Concretely, we build upon a key insight: in specific deployment settings, one can reveal carefully chosen non-sensitive operations (e.g. statistical calculations). This allows us to considerably improve the performance of privacy-preserving solutions without exposing the protocol to pervasive ML attacks. In more detail, we propose Soteria, a system for distributed privacy-preserving ML that leverages Trusted Execution Environments (e.g. Intel SGX) to run computations over sensitive information in isolated containers (enclaves) . Unlike previous work, where all ML-related computation is performed at trusted enclaves, we introduce a hybrid scheme, combining computation done inside and outside these enclaves. The experimental evaluation validates that our approach reduces the runtime of ML algorithms by up to 41% when compared to previous related work. Our protocol is accompanied by a security proof and a discussion regarding resilience against a wide spectrum of ML attacks.  \nINDEX TERMS Privacy-preserving, machine learning, distributed systems, apache spark, trusted execution environments, Intel SGX.  \nI. INTRODUCTION  \nThe ubiquitous environment provided by cloud computing providers offers a scalable, reliable, and performant environment to deploy compute-intensive Machine Learning (ML) workloads. However, many of these workloads operate over users’ sensitive information (e.g., medical records, financial information) . Regulations like HIPAA and GDPRenforce strong security policies when processing or storing sensitive data at untrusted third-party infrastructures [1],[2] . As such, outsourcing ML data storage and computation to third-party services leave users vulnerable to attacks that may compromise the integrity and confidentiality of their data [3], [4] . Indeed, the ML pipeline encompasses several  \nThe associate editor coordinating the review of this manuscript and  \napproving it for publication was Peter Langendoerfer  .  \nstages, both for model training and inference, in which users’data is known to be susceptible to different attacks such as adversarial attacks, model extraction, and inversion, and reconstruction attacks [5],[6],[7] .  \nRecent works have addressed these attacks with solutions based on homomorphic encryption or secure multi-party computation schemes [8],[9] . However, these cryptographic schemes impose a significant performance toll that restricts their applicability to practical scenarios [10] . To circumvent this performance penalty, another line of research is that of exploring hardware technologies enabling Trusted Execution Environments (TEEs), such as Intel SGX [11] . These technologies allow the execution of code within isolated processing environments (i.e., enclaves) where data can be securely handled in its original form (i.e., plaintext) at untrusted servers.  \nVOL","cbCaiaJLKkFCreKS","https://ap.wps.com/l/cbCaiaJLKkFCreKS","pdf",2670713,1,24,"English","en",105,"# Abstract\n# Introduction\n## Security and performance challenges in cloud ML\n## Limits of encryption-based approaches\n## Trusted Execution Environments (TEEs) and enclave performance trade-offs\n## Core challenge and proposed hybrid direction\n# System overview (Soteria)\n# Experimental evaluation\n# Security analysis","[{\"question\":\"Why do third-party cloud ML services require both security guarantees and performance considerations?\",\"answer\":\"Because training and inference are executed on sensitive data, strong confidentiality and integrity protections are needed. At the same time, security mechanisms can introduce substantial runtime overhead that reduces practical usability.\"},{\"question\":\"What is the main idea behind improving privacy-preserving ML performance in this work?\",\"answer\":\"Run most computations outside TEEs while revealing only carefully chosen non-sensitive operations (e.g., certain statistical calculations). This reduces enclave computational and I/O load without enabling pervasive ML attacks.\"},{\"question\":\"How does the Soteria approach differ from prior TEE-based privacy-preserving ML methods?\",\"answer\":\"Previous work typically performs the entire ML workload inside enclaves. Soteria uses a hybrid scheme that combines computation inside and outside the enclaves while still leveraging TEEs such as Intel SGX.\"}]","Privacy-Preserving Machine Learning on Apache Spark - Soteria Hybrid Trusted Execution Environments Scheme | PDF",1785813353,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"privacy-preserving-machine-learning-on-apache-spark-soteria-hybrid-trusted-execution-environments-scheme","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/privacy-preserving-machine-learning-on-apache-spark-soteria-hybrid-trusted-execution-environments-scheme/122857/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do third-party cloud ML services require both security guarantees and performance considerations?","Question",{"text":75,"@type":76},"Because training and inference are executed on sensitive data, strong confidentiality and integrity protections are needed. At the same time, security mechanisms can introduce substantial runtime overhead that reduces practical usability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main idea behind improving privacy-preserving ML performance in this work?",{"text":80,"@type":76},"Run most computations outside TEEs while revealing only carefully chosen non-sensitive operations (e.g., certain statistical calculations). This reduces enclave computational and I/O load without enabling pervasive ML attacks.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the Soteria approach differ from prior TEE-based privacy-preserving ML methods?",{"text":84,"@type":76},"Previous work typically performs the entire ML workload inside enclaves. Soteria uses a hybrid scheme that combines computation inside and outside the enclaves while still leveraging TEEs such as Intel SGX.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]