[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83861-en":3,"doc-seo-83861-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83861,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Towards Practical Human-Level Gaze Target Estimation","Gaze target estimation predicts where a person is looking in a scene, combining global scene semantics with precise spatial reasoning from human appearance cues such as pose and eye orientation. Existing vision models lag human performance, limiting use in interactive systems and robotics. The document introduces PaGE (Practical Gaze Estimator), which explicitly models scene–head feature interactions via a Scene-head Interaction Module and trains using a two-stage recipe plus token-level feature distillation on large unlabeled data.","arXiv :2607 .04860v 1 [ cs .CV] 6 Jul 2026  \nPAGE: TOWARDS PRACTICAL HUMAN-LEVEL GAZE TARGET ESTIMATION  \nZhoutong Ye 1∗ Chengwen Zhang 1∗ Zhaibin Cui 1 Mingze Sun 1  \nJiaqi Liu 1 Xiangwu Li2 Qingyang Wan 1 Chang Liu 1  \nXutong Wang 1 Huan-ang Gao 1 Yu Mei 1 Chun Yu 1† Yuanchun Shi 1  \n1 Tsinghua University 2Jinan University †Corresponding Author ∗Equal Contribution  \n{yezt24, [zcw25](zcw25}@mails.tsinghua.edu.cn)[}](zcw25}@mails.tsinghua.edu.cn)[@mails.tsinghua.edu.cn](zcw25}@mails.tsinghua.edu.cn)  \nABSTRACT  \nGaze target estimation, the task of predicting where a person is looking in a scene, is crucial to understanding human attention and intent. It is a challenging task that combines high-level understanding of global scene semantics and precise spatial reasoning using human appearance (e.g. pose, eye orientation) . As a result, human-level performance remains elusive for existing models, limiting their practical application. To this end, we propose PaGE (Practical Gaze Estimator), a gaze estimation model that explicitly models the complex interaction between scene and head features. Using a PaGE model with a large ViT-H+ backbone as the teacher, we further distill student models with lighter backbones on a much larger and more diverse unlabeled dataset. The architectural improvements and novel training recipe allow PaGE to achieve state-of-the-art performance on several gaze estimation tasks, outperforming humans in 7 out of 9 metrics while reducing the human-AI gap by at least 60% in the remaining 2 . The distilled student models retain most of the teacher’s performance while being lightweight enough for practical deployment on robots and consumer devices. The code and model checkpoints are available at [https://PaGE-26.github.io](https://PaGE-26.github.io).  \n1 INTRODUCTION  \nGaze is one of the most important non-verbal social cues. It provides valuable insight into a person’s attention and intent, as well as the dynamics of social interactions. It is a key component of sociallyaware interactive systems like MLLM agents and robots. Humans perform accurate gaze following (i.e., identifying the gaze target of another person in a scene) naturally, yet it is challenging to replicate this capability with vision models. This can be attributed to the inherent complexity of gaze following—it requires a combination of scene understanding and accurate spatial reasoning using human appearance cues (e.g., pose, eye orientation) . Therefore, existing models perform substantially worse than humans, limiting their practical application in fields like HCI and robotics.  \nIn this work, we propose PaGE (Figure 1), the first gaze estimation model with human-level performance. On GazeFollow, VideoAttentionTarget and ChildPlay, PaGE outperforms humans on 7 out of 9 metrics while closing the current human-AI gap by at least 60% for the remaining two. Distilled versions of PaGE retain SOTA performance while being lightweight enough for real-time gaze following on robots and many consumer devices.  \nThe strong results come from a combination of (1) a novel model architecture designed to explicitly model feature interaction between the scene and head branches, and (2) improvements to the training recipe, an area underexplored in previous work. Specifically, we propose the Scene-head Interaction Module (SIM), a novel gaze decoder component that uses cross attention between the  \nscene and head branches to explicitly model inter-branch feature interaction in a ViT-native manner. This affords PaGE the spatial reasoning capability needed to pinpoint the gaze target. For training, we adopt a new two-stage approach. We first train the decoder only, with the backbone frozen. We then finetune the entire model, backbone included, to further adapt the model to gaze prediction tasks. To build strong lightweight models, we further propose a token-level feature distillation procedure that trains student models using a PaGE ViT-H+ teacher. To ensure that th","cbCaidaw80jgX14I","https://ap.wps.com/l/cbCaidaw80jgX14I","pdf",3564958,3,1,21,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What is the core challenge in gaze target estimation described in the document?\",\"answer\":\"It requires combining global scene understanding with accurate spatial reasoning using human appearance cues such as pose and eye orientation.\"},{\"question\":\"How does PaGE achieve practical human-level performance?\",\"answer\":\"PaGE models interaction between scene and head features using a Scene-head Interaction Module with cross attention, and uses a training strategy with two-stage decoder training followed by full-model fine-tuning.\"},{\"question\":\"What role does token-level feature distillation play?\",\"answer\":\"It trains lightweight student models from a larger ViT-H+ teacher and uses large-scale unlabeled image data to build generalizable gaze-estimation features despite limited labeled annotations.\"}]",1784191037,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-practical-human-level-gaze-target-estimation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/towards-practical-human-level-gaze-target-estimation/83861/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core challenge in gaze target estimation described in the document?","Question",{"text":75,"@type":76},"It requires combining global scene understanding with accurate spatial reasoning using human appearance cues such as pose and eye orientation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does PaGE achieve practical human-level performance?",{"text":80,"@type":76},"PaGE models interaction between scene and head features using a Scene-head Interaction Module with cross attention, and uses a training strategy with two-stage decoder training followed by full-model fine-tuning.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does token-level feature distillation play?",{"text":84,"@type":76},"It trains lightweight student models from a larger ViT-H+ teacher and uses large-scale unlabeled image data to build generalizable gaze-estimation features despite limited labeled annotations.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]