[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121901-en":3,"doc-seo-121901-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121901,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Reconsideration on evaluation of machine learning models in continuous monitoring using wearables","This paper examines the difficulties of evaluating machine learning models for continuous health monitoring with wearable devices beyond standard segment-level metrics. It highlights how real-world signal variability, evolving disease dynamics, user-specific characteristics, and frequent false notifications complicate reliable assessment and deployment. Using lessons from large-scale heart studies, it proposes a comprehensive, standardized evaluation guideline aimed at producing robust models that remain accurate under diverse conditions and support safe, effective translation at population scale.","Reconsideration on evaluation of machine learning models in continuous monitoring using wearables  \nCheng Ding1,2, Zhicheng Guo3, Cynthia Rudin3,4, Ran Xiao1, Fadi B Nahab5, Xiao Hu1,2,6  \n1Nell Hodgson Woodruff School of Nursing, Emory University, Atlanta, GA, USA  \n2Wallace H. Coulter Department of Biomedical Engineering, Georgia Institute of Technology, Atlanta, GA, USA  \n3Department of Electrical and Computer Engineering, Duke University, Durham, NC, USA  \n4Department of Computer Science, Duke University, Durham, NC, USA  \n5Department of Neurology, Emory University School of Medicine, Atlanta, GA, USA  \n6Department of Biomedical Informatics, Emory University School of Medicine, Atlanta, GA, USA  \nAbstract  \nThis paper explores the challenges in evaluating machine learning (ML) models for continuous health monitoring using wearable devices beyond conventional metrics. We state the complexities posed by real-world variability, disease dynamics, user-specific characteristics, and the prevalence of false notifications, necessitating novel evaluation strategies. Drawing insights from large-scale heart studies, the paper offers a comprehensive guideline for robust ML model evaluation on continuous health monitoring.  \n1. Introduction  \nThe widespread adoption of electronic wearable devices with built-in biosensors has enabled their deployment to millions of users for various applications, including health condition monitoring, such as atrial fibrillation (AF) detection, blood pressure estimation, viral infections and blood oxygen saturation measurement [1-4] . Especially with the utilization of photoplethysmography (PPG) signal, these devices have demonstrated significant potential in providing real-time insights into an individual's health status. PPG, due to its non-invasive nature and ease of integration into wearable technology, has become a cornerstone in modern health monitoring systems [5] . Analyzing wearable device signals often involves ML models of different complexities [6, 7] . In the model development phase, typically, continuous signals are cut into discrete segments, and the model's performance is evaluated at the segment level using conventional metrics such as accuracy, sensitivity, specificity, and F1 score [8] . However, relying solely on these conventional metrics at the segment level does not provide a holistic assessment and hurts both consumers by making it impossible to select optimal solution for their needs and innovators by failing to guide their effort towards true progresses. The complex nature of continuous health monitoring using wearable devices introduces unique challenges beyond conventional evaluation approaches’capabilities, as illustrated in Figure 1. Recognizing these challenges is imperative for imbuing continuous health monitoring applications with accurate and reliable ML models to ensure a successful translation of these models into everyday use by millions of people and fulfill the potential of this technology at scale. In the subsequent sections, we outline the challenges in evaluating ML models for continuous health monitoring using wearables, thoroughly review existing evaluation methods and metrics, and propose a standardized evaluation guideline.  \nFigure 1. Challenges in Evaluating ML Models for Continuous Health Monitoring Using  \nWearables  \n2. Challenges in Evaluating ML Models for Continuous Health Monitoring Using Wearables  \n2.1 Real-World Variability and Environmental Factors  \nThe quality of physiological signal can be impacted by many factors, including user activities, time of the day, ambient lighting conditions, electromagnetic interference, room temperature, humidity, and other user behaviors (e.g. , compliance, comfort fit, etc. ) [9] . These factors can lead to morphological changes in the collected signal, even with additional introduction of artifacts , affecting the ML model's performance and introducing uncertainty in predictions. For example, Fig. 2a shows the","cbCaimdba2L7BN9T","https://ap.wps.com/l/cbCaimdba2L7BN9T","pdf",703403,1,10,"English","en",105,"# Introduction\n## Real-World Variability and Environmental Factors\n## Dynamic Nature of Diseases\n## User-Specific Characteristics","[{\"question\":\"Why do conventional segment-level metrics fail for continuous wearable monitoring?\",\"answer\":\"Segment-level metrics like accuracy, sensitivity, specificity, and F1 score do not provide a holistic view. They can mislead consumer choice and do not properly guide innovators toward improvements for real continuous use.\"},{\"question\":\"What real-world factors can degrade physiological signal quality in wearables?\",\"answer\":\"Signal quality can be affected by user activities, time of day, ambient lighting, electromagnetic interference, room temperature, humidity, and user behavior such as compliance and comfort fit. These factors can introduce artifacts and morphological changes that increase uncertainty.\"},{\"question\":\"How can disease dynamics limit evaluation based on short time windows?\",\"answer\":\"Some diseases show intermittent or prolonged patterns that short snapshots may miss. Metrics based on short intervals may overlook clinically important measures such as AF burden and related risk-relevant nuances.\"}]","Reconsideration on evaluation of machine learning models in continuous monitoring using wearables | PDF",1785807647,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"reconsideration-on-evaluation-of-machine-learning-models-in-continuous-monitoring-using-wearables","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/reconsideration-on-evaluation-of-machine-learning-models-in-continuous-monitoring-using-wearables/121901/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do conventional segment-level metrics fail for continuous wearable monitoring?","Question",{"text":75,"@type":76},"Segment-level metrics like accuracy, sensitivity, specificity, and F1 score do not provide a holistic view. They can mislead consumer choice and do not properly guide innovators toward improvements for real continuous use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What real-world factors can degrade physiological signal quality in wearables?",{"text":80,"@type":76},"Signal quality can be affected by user activities, time of day, ambient lighting, electromagnetic interference, room temperature, humidity, and user behavior such as compliance and comfort fit. These factors can introduce artifacts and morphological changes that increase uncertainty.",{"name":82,"@type":73,"acceptedAnswer":83},"How can disease dynamics limit evaluation based on short time windows?",{"text":84,"@type":76},"Some diseases show intermittent or prolonged patterns that short snapshots may miss. Metrics based on short intervals may overlook clinically important measures such as AF burden and related risk-relevant nuances.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]