[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127694-en":3,"doc-seo-127694-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127694,962084925782,"Ava Thompson","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Deception and defense from machine learning to supply chains - Doctor of Philosophy thesis","Broad classes of modern cyberattacks rely on deceiving human victims, and text is pervasive across computational systems. The thesis analyzes techniques that attack text encodings to generate deceptive inputs for critical applications. It shows adversarial, human-indistinguishable perturbations that manipulate text-based machine learning pipelines, affecting systems such as machine translation and toxic content detection, alongside corresponding defenses. It then introduces adversarial search to trigger selective search-engine results, including a social-engineering pathway to disinformation. Finally, it presents Trojan Source attacks in many programming languages, analyzes coordinated disclosures, and proposes ABOM for detecting software supply-chain attacks from binary-embedded dependency metadata.","Deception and defense from machine learning to supply chains  \nNicholas David Boucher  \nUniversity of Cambridge Computer Laboratory Clare College  \nDecember 2023  \nThis dissertation is submitted for  \nthe degree of Doctor of Philosophy  \nDeclaration  \nThis dissertation is the result of my own work and includes nothing which is the outcome of work done in collaboration except where specifically indicated in the text. It is not substantially the same as any work that has already been submitted before for any degree or other qualification.  \nThis dissertation does not exceed the regulation length of 60 , 000 words, including tablesand footnotes.  \nDeception and defense from machine learning to supply chains  \nNicholas Boucher  \nAbstract  \nBroad classes of modern cyberattacks are dependent upon their ability to deceive human victims. Given the ubiquity of text across modern computational systems, we present and analyze a set of techniques that attack the encoding of text to produce deceptive inputs to critical systems. By targeting a core building block of modern systems, we can adversarially manipulate dependent applications ranging from natural language processing pipelines to search engines to code compilers. Left undefended, these vulnerabilities enable many ill effects including uncurtailed online hate speech, disinformation campaigns, and software supply chain attacks.  \nWe begin by generating adversarial examples for text-based machine learning systems. Due to the discrete nature of text, adversarial examples for text pipelines have traditionally involved conspicuous perturbations compared to the subtle changes of the more continuous visual and auditory domains. Instead, we propose imperceptible perturbations: techniques that manipulate text encodings without affecting the text in its rendered form. We use these techniques to craft the first set of adversarial examples for text-based machine learning systems that are human-indistinguishable from their unperturbed form, and demonstrate their efficacy against systems ranging from machine translation to toxic content detection. We also describe a set of defenses against these techniques.  \nNext, we propose a new attack setting which we call adversarial search. In this setting, an adversary seeks to manipulate the results of search engines to surface certain results only and consistently when a hidden trigger is detected. We accomplish this by applying the encoding techniques of imperceptible perturbations to both indexed content and queries in major search engines. We demonstrate that imperceptibly encoded triggers can be used to manipulate the results of current commercial search engines, and then describe a social engineering attack exploiting this vulnerability that can be used to power disinformation campaigns. Again, we describe a set of defenses against these techniques.  \nWe then look to compilers and propose a different set of text perturbations which can be used to craft deceptive source code. We exploit the bidirectional nature of modern text standards to embed directionality control characters into comments and string literals. These control characters allow attackers to shuffle the sequence of tokens rendered in  \nsource code, and in doing so to implement programs that appear to do one thing when rendered to human code reviewers, but to do something different from the perspective of the compiler. We dub this technique the Trojan Source attack, and demonstrate the vulnerability of C, C++, C\\#, JavaScript, Java, Rust, Go, Python, SQL, Bash, Assembly, and Solidity. We also explore the applicability of this attack technique to launching supply chain attacks, and propose defenses that can be used to mitigate this risk. We also describe and analyze a 99-day coordinated disclosure that yielded patches to dozens of market-leading compilers, code editors, and code repositories.  \nFinally, we propose a novel method of identifying software supply chain attacks that works not ","cbCaiipqX7IplevZ","https://ap.wps.com/l/cbCaiipqX7IplevZ","pdf",8705970,4,1,161,"English","en",105,"# Abstract\n## Adversarial examples for text-based machine learning\n## Adversarial search against search engines\n## Trojan Source attacks in compilers and languages\n## Detecting supply chain attacks with ABOM\n## Thesis overview and defenses","[{\"question\":\"How does the thesis connect text encodings to cyberattacks?\",\"answer\":\"It argues that weaknesses in a core building block—text encodings—allow adversarial manipulation of downstream systems. By crafting deceptive inputs, attackers can influence machine learning pipelines, search engines, and compilers.\"},{\"question\":\"What are imperceptible perturbations and what do they enable?\",\"answer\":\"Imperceptible perturbations manipulate text encodings without changing how the text renders to humans. The thesis uses them to create human-indistinguishable adversarial examples for text-based machine learning systems.\"},{\"question\":\"What is the ABOM method and how does it detect supply chain attacks?\",\"answer\":\"ABOM embeds dependency metadata by hashing each consumed source file into compiled binaries, recursively including downstream dependencies. It then enables detection of poisoned dependencies in downstream software by querying binaries for embedded indicators.\"}]","Deception and defense from machine learning to supply chains - Doctor of Philosophy thesis | PDF",1785940930,406,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"deception-and-defense-from-machine-learning-to-supply-chains-doctor-of-philosophy-thesis","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/deception-and-defense-from-machine-learning-to-supply-chains-doctor-of-philosophy-thesis/127694/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How does the thesis connect text encodings to cyberattacks?","Question",{"text":76,"@type":77},"It argues that weaknesses in a core building block—text encodings—allow adversarial manipulation of downstream systems. By crafting deceptive inputs, attackers can influence machine learning pipelines, search engines, and compilers.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What are imperceptible perturbations and what do they enable?",{"text":81,"@type":77},"Imperceptible perturbations manipulate text encodings without changing how the text renders to humans. The thesis uses them to create human-indistinguishable adversarial examples for text-based machine learning systems.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the ABOM method and how does it detect supply chain attacks?",{"text":85,"@type":77},"ABOM embeds dependency metadata by hashing each consumed source file into compiled binaries, recursively including downstream dependencies. It then enables detection of poisoned dependencies in downstream software by querying binaries for embedded indicators.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]