[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127880-en":3,"doc-seo-127880-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127880,2336474459895,"Aria","https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916",8,"Research & Report","Fairness in Machine Learning Methods for Surgical Skill Assessment","Surgical skill development underpins high-quality patient care and successful surgical outcomes, yet conventional assessment approaches such as direct observation and outcome-based metrics face subjectivity and poor scalability. Video-based assessment with machine learning offers a more objective, scalable alternative, but deployment is threatened by bias, including bias originating in data. Using a pairwise comparison framework, this thesis introduces granular rater bias of varying intensities and quantifies its impact. Biased test sets yield an average AUC drop of 8% to 17.1% per 10% increase in rater bias. It further simulates realistic rater ratings from IAT scores of 131 surgeons, evaluates training with biased versus unbiased datasets, and proposes a fairness-metric-driven pipeline that quantifies dataset bias and enables downstream correction via reweighting and related techniques.","FAIRNESS IN MACHINE LEARNING METHODS FOR SURGICAL SKILL ASSESSMENT  \nby  \nNanthini Narayanan  \nA thesis submitted to Johns Hopkins University in conformity with the requirements for the degree of Master of Science in Engineering  \nBaltimore, Maryland  \nMay 2024  \n© 2024 Nanthini Narayanan  \nAll Rights Reserved  \nAbstract  \nSurgical skill development is crucial for ensuring high-quality patient care and successful surgical outcomes. Traditional methods of surgical skill assessment, such as direct observation and outcome-based metrics, often suffer from subjectivity and scalability issues. Video-based assessments (VBA) using machine learning (ML) could be a promising alternative, offering the potential for more objective and scalable evaluations. However, a significant challenge in deploying ML models in this context is the potential for bias, which can stem from various sources, including the data. Such biases can skew results, leading to unfair assessments and potentially impacting surgeon careers and patient outcomes.  \nUtilizing a pairwise comparison framework, we introduce bias of various intensities at a granular level and quantify it. Our findings highlight a significant difference in model evaluation on biased test sets, with an average 8% to 17. 1% drop in Area Under the Curve (AUC) scores for each 10% increase in rater bias. We also simulate realistic ratings for a sample of raters based on a study’s data on IAT scores of 131 surgeons, ensuring that the simulated data reflects the variability and distribution seen in real-world data.  \nNext, we evaluate the performance of models trained on biased and unbiased data sets, demonstrating that models trained on unbiased data outperform biased models in this case. We also propose a pipeline that begins with the quantification of dataset bias, which could be used to train models that compensate for identified biases downstream through reweighting and other corrective techniques. Central to our method is the use of fairness metrics, such as the true positive rate (TPR) and false positive rate (FPR) for equalized  \nodds, to measure bias. These metrics are calculated by comparing rater labels against expert labels within strategically sampled subsets of the data.  \nOur work underscores the necessity of addressing implicit bias in training and test sets to ensure the fairness and reliability of automated surgical skill assessment models.  \nThesis Reader: Dr. S. Swaroop Vedula, M. B. B.S, Ph.D, M.P.H.  \nThesis Reader: Dr. Shameema Sikder, M.D.  \nThesis Reader: Dr. Vishal M. Patel, Ph.D.  \nAcknowledgements  \nI am deeply grateful to Dr. S. Swaroop Vedula for the opportunity to be a part of his talented and encouraging team. His guidance and encouragement have been invaluable throughout my research journey. Dr. Vedula's patience, clear explanations, and readiness to engage in discussions significantly contributed to my growth and the success of this project.  \nI am thankful to Dr. Shameema Sikder for her support, valuable feedback, and insights from a surgeon’s perspective that enriched our discussions and presentations. I also extend my gratitude to Dr. Vishal Patel for granting access to crucial lab resources and for his guidance.  \nI am grateful to Divyasree Sasi Kumar, whose work laid the foundation for my own , for her generous help. I also want to thank Jay Paranjape, whose suggestions and assistance helped advance my experiments.  \nFinally, I would like to thank my family and friends for their constant love and support.  \nTable of Contents  \nAbstract........................................................................................................................... ii  \nAcknowledgements....................................................................................................... iv  \nTable of Contents ........................................................................................................... v  \nList of Tables ....................................","cbCaitc8Jq1FI9ue","https://ap.wps.com/l/cbCaitc8Jq1FI9ue","pdf",4695160,1,48,"English","en",105,"# Abstract\n# Chapter 1. Introduction\n## Surgical Skill Assessment\n## Machine Learning Methods for Surgical Skill Assessment\n## Fairness in Machine Learning Methods\n## Implicit Bias\n## Implicit Bias in ML Methods for Surgical Skill Assessment\n# Chapter 2. Datasets\n# Chapter 3. Methods\n## Implicit bias in model evaluation\n### Biased Test Sets\n### Model Architecture\n### Model Evaluation\n### Sensitivity Analysis\n## Implicit bias in model training\n## Proposed Method\n### Quantifyi","[{\"question\":\"Why is fairness important in automated surgical skill assessment using machine learning?\",\"answer\":\"Bias can enter through data and distort evaluations, leading to unfair assessments that may affect surgeons’ careers and patient outcomes.\"},{\"question\":\"How does the thesis quantify bias in model evaluation?\",\"answer\":\"It uses a pairwise comparison framework to introduce rater bias at a granular level and measures changes in model performance on biased test sets.\"},{\"question\":\"What fairness-related metrics are central to the proposed method?\",\"answer\":\"The method uses fairness metrics such as true positive rate (TPR) and false positive rate (FPR) for equalized odds, comparing rater labels against expert labels in strategically sampled data subsets.\"}]","Fairness in Machine Learning Methods for Surgical Skill Assessment | PDF",1785942547,121,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"fairness-in-machine-learning-methods-for-surgical-skill-assessment","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fairness-in-machine-learning-methods-for-surgical-skill-assessment/127880/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is fairness important in automated surgical skill assessment using machine learning?","Question",{"text":76,"@type":77},"Bias can enter through data and distort evaluations, leading to unfair assessments that may affect surgeons’ careers and patient outcomes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis quantify bias in model evaluation?",{"text":81,"@type":77},"It uses a pairwise comparison framework to introduce rater bias at a granular level and measures changes in model performance on biased test sets.",{"name":83,"@type":74,"acceptedAnswer":84},"What fairness-related metrics are central to the proposed method?",{"text":85,"@type":77},"The method uses fairness metrics such as true positive rate (TPR) and false positive rate (FPR) for equalized odds, comparing rater labels against expert labels in strategically sampled data subsets.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]