[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119896-en":3,"doc-seo-119896-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119896,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Personalising lung cancer screening with machine learning","Personalised screening links repeated risk assessment to tailored management, yet large-scale delivery is complex. This dissertation develops machine learning approaches to simplify lung cancer risk assessment and to enable privacy-preserving analytics via synthetic data when patient records are unavailable. Parsimonious models predict five-year lung cancer risk, trained on UK Biobank and US NLT participants and externally validated in distinct ever-smokers. Ensemble models using age, smoking duration, and pack-years match or exceed key comparators, and synthetic UK Biobank data support prognostic model development while preserving statistical relationships.","Personalising lung cancer screening with machine learning  \nThomas Callender  \nA dissertation submitted in partial fulﬁllment  \nof the requirements for the degree of  \nDoctor of Philosophy  \nof  \nUniversity College London.  \nDepartment of Respiratory Medicine  \nUniversity College London  \nAugust 17, 2023  \n2  \nI, Thomas Callender, conﬁrm that the work presented in this thesis is my own. Where information has been derived from other sources, I conﬁrm that this has been indicated in the work.  \nAbstract  \nPersonalised screening is based on a straightforward concept: repeated risk assessment linked to tailored management. However, delivering such programmes at scale is complex. In this work, I aimed to contribute to two areas: the simpliﬁcation of risk assessment to facilitate the implementation of personalised screening for lung cancer; and, the use of synthetic data to support privacy-preserving analytics in the absence of access to patient records.  \nI ﬁrst present parsimonious machine learning models for lung cancer screening, demonstrating an approach that couples the performance of model-based risk prediction with the simplicity of risk-factor-based criteria. I trained models to predict the ﬁve-year risk of developing or dying from lung cancer using UK Biobank and US National Lung Screening Trial participants before external validation amongst temporally and geographically distinct ever-smokers in the US Prostate, Lung, Colorectal and Ovarian Screening trial. I found that three predictors – age, smoking duration, and pack-years – within an ensemble machine learning framework achieved or exceeded parity in discrimination, calibration, and net beneﬁt with comparators. Furthermore, I show that these models are more sensitive than risk-factor-based criteria, such as those currently recommended by the US Preventive Services Taskforce.  \nFor the implementation of more personalised healthcare, researchers and developers require ready access to high-quality datasets. As such data are sensitive, their use is subject to tight control, whilst the majority of data present in electronic records are not available for research use. Synthetic data are algorithmically generated but can maintain the statistical relationships present within an original dataset. In this work, I used explicitly privacypreserving generators to create synthetic versions of the UK Biobank before we performed exploratory data analysis and prognostic model development. Comparing results when using  \nAbstract 4  \nthe synthetic against the real datasets, we show the potential for synthetic data in facilitating prognostic modelling.  \nImpact Statement  \nLung cancer remains the most common cause of death from cancer worldwide. For those who are considered at high-risk of lung cancer, regular screening with low-dose CT scans has been shown to reduce lung-cancer deaths by 20% . Because of this the UK, like many other countries worldwide, is in the process of developing a national lung cancer screening programme. A key question is how to determine whether someone is at high risk. In other words, who should we screen?  \nIn this work, I’ve used ensemble machine learning to develop new prediction models for lung cancer that perform as well as those in use whilst needing only three predictors, one-quarter of those required by comparators. These new models could ease lung cancer risk assessment. In turn, this could lead to alternative ways of implementing a national screening programme that might increase uptake and therefore the beneﬁts of the programme.  \nThe translational aspect of this work could also extend beyond lung cancer. Screening and disease prevention are increasingly personalised, based on a repeating cycle of risk assessment and tailored management. Although an intuitive concept, putting this into practice is remarkably difﬁcult at scale because most of the predictors we need are not present in electronic health records, and if they are present, may not be accura","cbCaisf5tWO3AqSs","https://ap.wps.com/l/cbCaisf5tWO3AqSs","pdf",9734367,1,210,"English","en",105,"# Abstract\n# Impact Statement\n# Acknowledgements","[{\"question\":\"What problem does this dissertation address in personalised lung cancer screening?\",\"answer\":\"It addresses the difficulty of delivering personalised screening at scale by simplifying risk assessment and by enabling analytics when access to patient records is restricted.\"},{\"question\":\"Which lung-cancer risk predictors are highlighted in the proposed models?\",\"answer\":\"The work emphasizes age, smoking duration, and pack-years, using an ensemble machine learning framework to predict five-year risk.\"},{\"question\":\"How does synthetic data contribute to the research approach?\",\"answer\":\"Explicitly privacy-preserving synthetic generators create synthetic versions of UK Biobank that maintain statistical relationships, supporting exploratory analysis and prognostic model development without using real patient records.\"}]","Personalising lung cancer screening with machine learning | PDF",1785726887,529,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"personalising-lung-cancer-screening-with-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/personalising-lung-cancer-screening-with-machine-learning/119896/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this dissertation address in personalised lung cancer screening?","Question",{"text":75,"@type":76},"It addresses the difficulty of delivering personalised screening at scale by simplifying risk assessment and by enabling analytics when access to patient records is restricted.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which lung-cancer risk predictors are highlighted in the proposed models?",{"text":80,"@type":76},"The work emphasizes age, smoking duration, and pack-years, using an ensemble machine learning framework to predict five-year risk.",{"name":82,"@type":73,"acceptedAnswer":83},"How does synthetic data contribute to the research approach?",{"text":84,"@type":76},"Explicitly privacy-preserving synthetic generators create synthetic versions of UK Biobank that maintain statistical relationships, supporting exploratory analysis and prognostic model development without using real patient records.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]