[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122967-en":3,"doc-seo-122967-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122967,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","EFFICIENT OBSERVATION TIME WINDOW SEGMENTATION FOR ADMINISTRATIVE DATA MACHINE LEARNING - A PREPRINT","Machine learning models gain value from temporal trends in time-stamped administrative data, which are commonly encoded by splitting an observation window into time bins. Fine-grained binning per feature can improve performance, but it expands the hyperparameter search space exponentially. The paper proposes TAIB (Binning for Time Series Analysis), a computationally efficient approach using dynamic time warping to rank feature subsets by their benefit from bin-size tuning. Evaluations on hospital and housing/homelessness datasets show faster training and equal or improved predictive performance.","arXiv :2401 . 16537v2 [ cs .LG] 12 Mar 2024  \nEFFICIENT OBSERVATION TIME WINDOW SEGMENTATION FOR ADMINISTRATIVE DATA MACHINE LEARNING  \nA PREPRINT  \nMusa Taib and Geoffrey G. Messier  \nUniversity of Calgary  \n2500 University Dr. NW, Calgary, AB, Canada, T2N 1N4  \ngmessier@ucalgary.ca  \nABSTRACT  \nMachine learning models bene􀀂t when allowed to learn from temporal trends in time-stamped administrative data. These trends can be represented by dividing a model’s observation window into time segments or bins. Model training time and performance can be improved by representing each feature with a different time resolution. However, this causes the time bin size hyperparameter search space to grow exponentially with the number of features. The contribution of this paper is to propose a computationally ef􀀂cient time series analysis to investigate binning (TAIB) technique that determines which subset of data features bene􀀂t the most from time bin size hyperparameter tuning.  \nThis technique is demonstrated using hospital and housing/homelessness administrative data sets.  \nThe results show that TAIB leads to models that are not only more ef􀀂cient to train but can perform better than models that default to representing all features with the same time bin size.  \nKeywords administrative data, time window segmentation, machine learning, hospital, homelessness  \n1 Introduction  \nUtilizing data features from administrative data to predict outcomes for individuals or groups is an exciting application of machine learning. Administrative data typically consist of records collected while delivering services that can include the justice system, 􀀂ling taxes, progressing through the education system or accessing housing and homelessness services [22, 38, 6, 16] . However, with the extensive use of electronic medical records (EMRs), it is arguable that healthcare is the one of the most important sectors to make use of administrative data [24] .  \nSince most administrative data entries are time stamped, administrative data are fundamentally time series data. The time dimension carries important trend information for a machine learning model like whether a patient’s condition in hospital is improving or getting worse with time. The length of time between events is also important since thereis clearly a difference between someone who accesses an emergency housing shelter seven times in one week versus seven times in one year. The importance of the time dimension is not lost on most machine learning researchers. For example, the excellent survey by Morid, [et. al](et. al). [27] summarizes over 76 papers applying machine learning to medical administrative data and distinguishes them in part by how they account for the temporal nature of that data.  \nWhile a popular approach is to simply use temporal information to organize data features into an ordered sequence, understanding when an event does not occur can be as valuable as representing as when it does [20] . As a result, a large number of studies represent temporal administrative data using the temporal matrix approach [27] . In a temporal matrix, columns correspond to speci􀀂c events or data features and rows correspond to regularly spaced time intervals within the observation window. These time intervals are referred to as time bins.  \nTime bin duration or “size” is an important temporal matrix parameter. While the simplistic approach is to use the same bin size for all data features, it has been demonstrated that performance and model complexity can be improved when representing different features using different time resolutions [36, 1] . However, the challenge is that allowing each feature to have its own bin size means that the model’s hyperparameter search space will grow exponentially with the number of data features.  \nThe contribution of this paper is to present a pre-processing technique that determines which data features bene􀀂t the most from time bin size tuning and which features can be ","cbCaijqwz46KAZkS","https://ap.wps.com/l/cbCaijqwz46KAZkS","pdf",467213,1,14,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does TAIB address in machine learning on administrative time-stamped data?\",\"answer\":\"TAIB addresses the challenge that using different time bin sizes per feature can improve model accuracy but makes the time bin size hyperparameter search space grow exponentially with the number of features.\"},{\"question\":\"How does TAIB decide which features should receive time bin size tuning?\",\"answer\":\"TAIB ranks features using a metric derived from dynamic time warping, estimating how much discriminative power improves as features are represented at higher time resolution. It then tunes a small top subset while representing remaining features with a single bin size.\"},{\"question\":\"What datasets and outcomes are used to evaluate the approach?\",\"answer\":\"The technique is demonstrated on hospital and housing/homelessness administrative datasets, where TAIB produces models that train more efficiently and can achieve performance equal to or better than using the same bin size for all features.\"}]","EFFICIENT OBSERVATION TIME WINDOW SEGMENTATION FOR ADMINISTRATIVE DATA MACHINE LEARNING - A PREPRINT | PDF",1785813939,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"efficient-observation-time-window-segmentation-for-administrative-data-machine-learning-a-preprint","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/efficient-observation-time-window-segmentation-for-administrative-data-machine-learning-a-preprint/122967/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TAIB address in machine learning on administrative time-stamped data?","Question",{"text":75,"@type":76},"TAIB addresses the challenge that using different time bin sizes per feature can improve model accuracy but makes the time bin size hyperparameter search space grow exponentially with the number of features.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TAIB decide which features should receive time bin size tuning?",{"text":80,"@type":76},"TAIB ranks features using a metric derived from dynamic time warping, estimating how much discriminative power improves as features are represented at higher time resolution. It then tunes a small top subset while representing remaining features with a single bin size.",{"name":82,"@type":73,"acceptedAnswer":83},"What datasets and outcomes are used to evaluate the approach?",{"text":84,"@type":76},"The technique is demonstrated on hospital and housing/homelessness administrative datasets, where TAIB produces models that train more efficiently and can achieve performance equal to or better than using the same bin size for all features.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]