[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84293-en":3,"doc-seo-84293-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84293,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies","Autism spectrum disorder (ASD) affects over 75 million individuals worldwide, yet scalable computational methods for remote behavioral screening remain limited. This study evaluates automated detection of autism-related self-stimulatory hand behaviors from video by optimizing both sequence-based neural architectures and temporal sampling rates. Long short-term memory (LSTM) and gated recurrent unit (GRU) models outperform prior CNN baselines, reaching peak accuracies near 97.5–98.75% at a 15-frame interval. The work further assesses ten augmentation strategies for I3D transfer learning and tests personalized per-subject modeling under temporally split segments.","arXiv :2607 .07957v 1 [ cs .AI] 8 Jul 2026  \nEvaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand  \nIdiosyncrasies  \nRaunak Mondal 1 and Peter Washington2  \n1 School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA  \n2 Department of Information and Computer Sciences,  \nUniversity of Hawai‘i at M¯anoa, Honolulu, HI, USA  \nAbstract  \nAutism spectrum disorder (ASD) affects over 75 million individuals worldwide, yet scalable computational methods for remote behavioral screening remain limited. This study addresses two complementary challenges in automated detection of autism-related self-stimulatory behaviors from video: (1) identifying the optimal sequence-based neural network architecture and temporal sampling rate, and (2) characterizing data augmentation strategies for training on small behavioral datasets. For the first objective, long short-term memory (LSTM) and gated recurrent unit (GRU) models were trained on pose-derived features from the Self-Stimulatory Behavior Diagnosis (SSBD) dataset at frame sampling intervals of 1, 5, 15, 30, 45, and 90 frames. Both architectures exceeded prior convolutional neural network (CNN) baselines (62–76% accuracy), with peak accuracies of 97.5%(LSTM) and 98.75%(GRU) at a sampling interval of every  \n15 frames. For the second objective, ten data augmentation strategies were applied to an I3D transfer learning pipeline, with an ablation study quantifying the marginal contribution of each technique. Horizontal flip achieved the highest standalone accuracy (48.78%), while exclusion of upsampling from the augmentation pipeline produced the largest performance degradation, indicating its necessity for augmentation applied to complex behavioral video. A personalized machine learning approach, in which per-subject models were trained and tested on temporally split segments of each video, produced consistent predictions (mean loss 1.84, SD 0.79) . These results provide practitioners with concrete guidance on architecture selection, sampling rate, and augmentation strategy for video-based behavioral classification in data-scarce clinical domains.  \n1 Introduction  \nAutism spectrum disorder (ASD) is a neurodevelopmental condition characterized by differences in social communication and the presence of restricted, repetitive behaviors [1] . Current estimates place worldwide prevalence at approximately 75 million individuals, with reported rates in the United States increasing 241% since 2000 [2] . Despite this high prevalence, diagnosis remains labor-intensive: clinicians rely on structured behavioral observation instruments such as the Autism Diagnostic Observation Schedule (ADOS-2), which require trained administrators and multi-visit assessment pipelines [3] . These constraints create bottlenecks that delay diagnosis, particularly in low-resource settings where specialist access is limited.  \nAutomated recognition of self-stimulatory (“stimming”) behaviors from video offers one path toward scalable screening. Behaviors such as hand flapping, head banging, and spinning are among the motor stereotypies commonly observed in ASD and can in principle be detected from consumergrade video recordings [4] . However, computational approaches face two interrelated challenges.  \nFirst, publicly available datasets of labeled stimming behaviors are small. The Self-Stimulatory Behavior Diagnosis (SSBD) dataset [4], one of the few purpose-built resources, contains only 75 videos across three behavior categories. Second, prior work has relied heavily on convolutional neural network (CNN) models, which achieve state-of-the-art accuracies of only 62–76% on these tasks [5] .  \nSequential architectures such as long short-term memory (LSTM) networks and gated recurrent units (GRUs) are well-suited to temporal pattern recognition and have shown strong performance in adjacent activity recognition domains [6, 7] . Yet their application to autism-s","cbCaigbXK2Ys69kg","https://ap.wps.com/l/cbCaigbXK2Ys69kg","pdf",341856,4,1,15,"English","en",105,"# Abstract\n# Introduction\n## Dataset constraints and baseline performance\n## Sequential models and temporal granularity\n## Transfer learning and augmentation strategies\n## Experimental contributions","[{\"question\":\"What main problems does the study address for video-based autism screening?\",\"answer\":\"It tackles two challenges: selecting an optimal sequence-based model architecture with the right temporal sampling rate, and determining effective data augmentation strategies for training on small behavioral datasets.\"},{\"question\":\"Which models are evaluated, and what temporal sampling interval performs best?\",\"answer\":\"The study trains LSTM and GRU models on pose-derived features and tests multiple frame sampling intervals. Peak performance occurs at a sampling interval of every 15 frames, reaching about 97.5% for LSTM and about 98.75% for GRU.\"},{\"question\":\"How do augmentation strategies affect performance in the I3D transfer learning pipeline?\",\"answer\":\"Ten augmentation methods are tested with an ablation study to measure each technique’s marginal contribution. Horizontal flip gives the best standalone accuracy, while removing upsampling causes the largest performance drop, indicating upsampling is critical for complex behavioral video.\"}]",1784194617,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-the-effect-of-frame-rate-in-sequence-based-classification-of-autism-related-self-stimulatory-hand-idiosyncrasies","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/evaluating-the-effect-of-frame-rate-in-sequence-based-classification-of-autism-related-self-stimulatory-hand-idiosyncrasies/84293/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What main problems does the study address for video-based autism screening?","Question",{"text":75,"@type":76},"It tackles two challenges: selecting an optimal sequence-based model architecture with the right temporal sampling rate, and determining effective data augmentation strategies for training on small behavioral datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which models are evaluated, and what temporal sampling interval performs best?",{"text":80,"@type":76},"The study trains LSTM and GRU models on pose-derived features and tests multiple frame sampling intervals. Peak performance occurs at a sampling interval of every 15 frames, reaching about 97.5% for LSTM and about 98.75% for GRU.",{"name":82,"@type":73,"acceptedAnswer":83},"How do augmentation strategies affect performance in the I3D transfer learning pipeline?",{"text":84,"@type":76},"Ten augmentation methods are tested with an ablation study to measure each technique’s marginal contribution. Horizontal flip gives the best standalone accuracy, while removing upsampling causes the largest performance drop, indicating upsampling is critical for complex behavioral video.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]