[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85181-en":3,"doc-seo-85181-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85181,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","TextGaze Gaze Target Estimation with Textual Scene Cues","Gaze target estimation predicts where a person’s gaze points within a scene and whether the target lies inside the image. Conventional multi-branch designs rely on extra supervision and heavy annotations, while streamlined approaches overfit low-level visual saliency and misalign attention with true gaze targets. TextGaze addresses this conflict with a unified cross-modal architecture that uses a Large Vision-Language Model as scalable semantic guidance. A frozen visual encoder extracts features, an LVLM provides gaze-aligned textual cues, and a transformer fusion module with hierarchical text supervision preserves task semantics. Lightweight decoding jointly predicts gaze heatmaps and in-/out-of-frame status, achieving competitive results on four mainstream datasets with strong cross-dataset generalisation without additional fine-tuning.","TextGaze: Prompting Gaze Target Estimation with Textual Scene Cues  \nJunhui She 1 ,3 , Fei Wang2 ,3 , ∗ , Kun Li5 , Yiqi Nie3 ,4 , Yuxin Liu3 ,4 , Zhangling Duan3 and Xun Yang 1 ,∗  \n1University of Science and Technology of China  \n2Hefei University of Technology  \n3Institute of Artificial Intelligence, Hefei Comprehensive National Science Center  \n4 Anhui University  \n5 United Arab Emirates University  \n{shejunhui323, ifei17.hfut, kunli.hfut, nieyiqi5, [yuxinliu221](yuxinliu221}@gmail.com)[}](yuxinliu221}@gmail.com)[@gmail.com](yuxinliu221}@gmail.com),  \n[duanzl1024@iai.ustc.edu.cn](duanzl1024@iai.ustc.edu.cn), [xyang21@ustc.edu.cn](xyang21@ustc.edu.cn)  \narXiv :2607 . 10130v1 [ cs .CV] 11 Jul 2026  \nAbstract  \nGaze target estimation aims to infer the position of a person’s gaze within a scene. Within mainstream design logic, multi-branch methods require extra supervision and annotations, while streamlined designs prioritize low-level visual saliency over true gaze intent. The former leads to a high annotation burden and hinders domain transfer, whereas the latter causes misalignment between predicted attention and actual gaze targets. To address this issue, we propose TextGaze, a unified cross-modal architecture that leverages a Large Vision-Language Model (LVLM) as scalable semantic guidance to balance the two design paradigms. The model extracts visual features from a frozen encoder and utilizes an LVLM to obtain gaze-aligned textual cues. We design a transformer-based fusion module with hierarchical text supervision to preserve task semantics. Lightweight decoding heads enable the joint prediction of gaze heatmaps and in-/outof-frame status. We evaluate our method on four mainstream datasets, and the results show competitive performance across key metrics with robust cross-dataset generalisation without extra finetuning. Overall, we provide a streamlined alternative to traditional designs and highlight the potential of LVLMs as accessible auxiliary guidance  \nfor gaze estimation. All contents are available at: [https://github.com/idremo/TextGaze-IJCAI2026](https://github.com/idremo/TextGaze-IJCAI2026) .  \n1 Introduction  \nGaze target estimation Recasens et al. [2017]; Chong et al.[2018]; Miao et al. [2023]; Ryan et al. [2025]; Liu et al.[2024] is a task in computer vision that aims to predict a person’s gaze direction in a scene or determine if the gaze target lies within the image. As human gaze is a fundamental nonverbal cue reflecting attention allocation, intent, and cognitive  \n∗ Corresponding authors  \nFigure 1: Pipelines for gaze target estimation. (a) Multi-branch methods Chen et al. [2021] suffer from slow convergence due to increased structural complexity. (b) Prompt-tuning method Ryan et al.[2025] encounter semantic defocus as they struggle to isolate taskrelevant saliency via positional prompts. By contrast, our method (c) leverages accessible scene-level text to achieve precise semantic focus on the gaze target.  \nstate Eckstein et al. [2017], it is essential for human behavior understanding Li et al. [2023b, 2025, 2026]; Wang et al.[2024b], with broad applications in human-computer interaction Katsini et al. [2020]; Wang et al. [2026]; Chen et al.[2025], extended reality Sitzmann et al. [2018], and education Wang et al. [2025] .  \nIn multi-branch methods, the classical architecture is the dual-stream appearance-based design Chong et al. [2018]; Lian et al. [2018]; Chong et al. [2020]; Wang et al. [2024a] . In these designs, the head stream localizes the subject, and the scene stream encodes contextual visual information. Because purely visual information is insufficient to enhance effects,  \nthis paradigm has evolved into multi-branch frameworks that incorporate auxiliary cues. These include head features Chenet al. [2021], scene context Saran et al. [2018], depth Bao et al. [2022], and pose Gupta et al. [2022] .  \nAlthough these designs enrich visual representations, they introduced the slow convergence pro","cbCaivzC3DbW2Oo9","https://ap.wps.com/l/cbCaivzC3DbW2Oo9","pdf",1911452,2,1,10,"English","en",105,"# Introduction\n## Background and limitations\n## Multi-branch designs\n## Prompt-tuning with foundation models\n## TextGaze approach and motivation","[{\"question\":\"What problem does TextGaze address in gaze target estimation?\",\"answer\":\"TextGaze targets the mismatch caused by models that either require excessive multi-branch supervision or rely on low-level visual saliency that fails to reflect true gaze intent and semantic target selection.\"},{\"question\":\"How does TextGaze use a Large Vision-Language Model (LVLM)?\",\"answer\":\"It extracts visual features with a frozen encoder and then uses the LVLM to generate gaze-aligned textual cues that provide scalable semantic guidance for the estimation task.\"},{\"question\":\"What does TextGaze predict during inference?\",\"answer\":\"It uses lightweight decoding heads to jointly predict gaze heatmaps and whether the gaze target is in-frame or out-of-frame.\"}]",1784201584,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"textgaze-gaze-target-estimation-with-textual-scene-cues","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/textgaze-gaze-target-estimation-with-textual-scene-cues/85181/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TextGaze address in gaze target estimation?","Question",{"text":75,"@type":76},"TextGaze targets the mismatch caused by models that either require excessive multi-branch supervision or rely on low-level visual saliency that fails to reflect true gaze intent and semantic target selection.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TextGaze use a Large Vision-Language Model (LVLM)?",{"text":80,"@type":76},"It extracts visual features with a frozen encoder and then uses the LVLM to generate gaze-aligned textual cues that provide scalable semantic guidance for the estimation task.",{"name":82,"@type":73,"acceptedAnswer":83},"What does TextGaze predict during inference?",{"text":84,"@type":76},"It uses lightweight decoding heads to jointly predict gaze heatmaps and whether the gaze target is in-frame or out-of-frame.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]