[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119419-en":3,"doc-seo-119419-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119419,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Synergising Program Analysis and Machine Learning for Program Repair","Automated program repair (APR) improves software quality by fixing bugs automatically, yet it is constrained by program complexity. The huge space of possible patches makes exhaustive search impractical, while generated fixes may be incorrect and only overfit test cases. This work addresses the gap by combining code structure signals: program analysis captures execution semantics and semantic behavior, while machine learning leverages large data to interpret natural-language elements such as identifiers and comments. The thesis presents four complementary research contributions for semantic repair, search-space reduction, probabilistic state approximation, and fact selection for LLM-based APR prompts.","Synergising Program Analysis and Machine Learning for Program  \nRepair  \nNikhil Parasaram  \nA dissertation submitted in partial fulfillment of the requirements for the degree of  \nDoctor of Philosophy  \nof  \nUniversity College London.  \nDepartment of Computer Science  \nUniversity College London  \nOctober 21, 2024  \n2  \nI, Nikhil Parasaram, confirm that the work presented in this thesis is my own. Where information has been derived from other sources, I confirm that this has been indicated in the work. The following chapters are based on the publications listed below:  \n• Chapter 3: Parasaram, Nikhil, Earl T. Barr, and Sergey Mechtaev. \"Trident: Controlling side effects in automated program repair.\" IEEE Transactions on Software Engineering 48.12 (2021) .  \n• Chapter 4: Parasaram, Nikhil, Earl T. Barr, and Sergey Mechtaev. \"Rete: Learning namespace representation for program repair.\" 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) . IEEE, 2023 .  \n• Chapter 5: Parasaram, Nikhil, Earl T. Barr, Sergey Mechtaev, and Marcel Böhme. \"Precise Data-Driven Approximation for Program Analysis via Fuzzing.\" In 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE), pp. 611-623. IEEE, 2023 .  \n• Chapter 6: Parasaram, Nikhil, Huijie Yan, Boyu Yang, Zineb Flahy, Abriele Qudsi, Damian Ziaber, Earl T. Barr, and Sergey Mechtaev. \"The Fact Selection Problem in LLM-Based Program Repair.\" arXiv preprint arXiv:2404.05520  \n(2024) . Submitted to a Conference, currently under review.  \nAbstract  \nAutomated program repair (APR) enhances software quality by fixing bugs automatically, but it faces challenges due to software complexity. The vast number of possible patches makes exhaustive search impractical, and identifying correct patches is difficult since tools may generate incorrect fixes that overfit test cases. Mitigating these challenges involves leveraging the structure of code, which consists of a formal channel (execution semantics) and a natural language channel (comments, variable names) . Machine learning excels at interpreting the natural language channel using large datasets but struggles with generating semantically correct patches. Conversely, program analysis provides detailed insights into program semantics. Combining program analysis with machine learning can address these challenges, using program analysis for execution specifics and machine learning for natural code aspects, like identifiers and comments. This thesis consists of four different works:  \n• Chapter 3 advances semantic repair by synthesizing patches with side effects, employing symbolic execution with state merging and effective patch prioritization to repair bugs in open-source projects.  \n• Chapter 4 reduces the search space by utilizing neural networks to learn variable information, ranking variables and patch templates to improve accuracy and reduce test overfitting. This enhances existing approaches, allowing them to repair previously unfixable bugs by leveraging program namespace information.  \n• Chapter 5 leverages abstract interpretation and fuzzing to probabilistically approximate reachable program states, focusing on high-probability states.  \nAbstract 4  \nPSP boosts performance in abstract interpretation, symbolic execution, and patch prioritization, benefiting strategies discussed in Chapter 3 and Chapter 4 .  \n• Chapter 6 addresses the challenge of identifying the most relevant facts, such as test errors and angelic values, for constructing effective prompts for LLM based APR. Extracted through program analysis, these facts build prompts whose effectiveness varies across different bugs. We develop a strategy to select facts tailored to each specific bug, significantly enhancing the effectiveness of LLMs in APR.  \nImpact Statement  \nWith the advent of Generative AI, much research has focused on directly applying machine learning to various tasks, including automated program repair. While machine learning","cbCaijXIwaESqTZR","https://ap.wps.com/l/cbCaijXIwaESqTZR","pdf",1908520,1,202,"English","en",105,"# Abstract\n## Chapter Contributions\n## Impact Statement\n## Published Work and Awards\n## Key Techniques by Chapter","[{\"question\":\"What are the main challenges that automated program repair (APR) faces?\",\"answer\":\"APR must handle the extremely large patch search space, making exhaustive exploration infeasible. It also struggles to distinguish genuinely correct fixes from incorrect ones that overfit test cases.\"},{\"question\":\"How does the thesis combine program analysis and machine learning?\",\"answer\":\"Program analysis is used to capture execution semantics and detailed program behavior, while machine learning is used to interpret natural-language aspects of code, such as comments and identifiers. Together, they aim to generate semantically correct patches.\"},{\"question\":\"What does each chapter contribute to the overall APR goal?\",\"answer\":\"Chapter 3 synthesizes patches with side effects using symbolic execution and prioritization. Chapter 4 learns variable and template information to reduce search and improve accuracy. Chapter 5 uses abstract interpretation and fuzzing to approximate reachable states probabilistically. Chapter 6 selects the most relevant facts to craft effective prompts for LLM-based APR.\"}]","Synergising Program Analysis and Machine Learning for Program Repair | PDF",1785724202,509,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"synergising-program-analysis-and-machine-learning-for-program-repair","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/synergising-program-analysis-and-machine-learning-for-program-repair/119419/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the main challenges that automated program repair (APR) faces?","Question",{"text":75,"@type":76},"APR must handle the extremely large patch search space, making exhaustive exploration infeasible. It also struggles to distinguish genuinely correct fixes from incorrect ones that overfit test cases.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis combine program analysis and machine learning?",{"text":80,"@type":76},"Program analysis is used to capture execution semantics and detailed program behavior, while machine learning is used to interpret natural-language aspects of code, such as comments and identifiers. Together, they aim to generate semantically correct patches.",{"name":82,"@type":73,"acceptedAnswer":83},"What does each chapter contribute to the overall APR goal?",{"text":84,"@type":76},"Chapter 3 synthesizes patches with side effects using symbolic execution and prioritization. Chapter 4 learns variable and template information to reduce search and improve accuracy. Chapter 5 uses abstract interpretation and fuzzing to approximate reachable states probabilistically. Chapter 6 selects the most relevant facts to craft effective prompts for LLM-based APR.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]