[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85896-en":3,"doc-seo-85896-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85896,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","GigaChat Audio Time-aware Large Audio Language Model","Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. This work introduces a time-aware audio language model that answers questions with explicit timestamps for up to 120 minutes of input. The method interleaves periodic time markers with continuous audio tokens and trains via large-scale synthetic supervision from a cascaded pipeline. Experiments validate temporal-grounding accuracy across short and long benchmarks, and enable time-anchored fragment descriptions and summaries, supported by extensive ablations. Open model weights and datasets are released for further research.","GigaChat Audio: Time-aware Large Audio Language Model  \nAleksandr Kutsakov 1, Mariia Sadovina 1 , Georgii Gospodinov 1 , Alexandr Maximenko 1, Oleg Kutuzov 1, Pavel Bogomolov 1, Fyodor Minkin 1  \n1 SaluteDevices, Russia  \n{askutsakov, sadovinama, georgygospodinov, ae .maximenko, olegkutuzov01, bobrosoft98,  \n[minkin.fyodor](minkin.fyodor}@gmail.com)[}](minkin.fyodor}@gmail.com)[@gmail.com](minkin.fyodor}@gmail.com)  \narXiv :2607 . 10387v1 [ ee ss .AS] 11 Jul 2026  \nAbstract  \nTemporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves periodic time markers with continuous audio tokens using large-scale synthetic supervision from a cascaded pipeline. Our model achieves strong temporal-grounding accuracy on short and long benchmarks and supports time-anchored fragment descriptions and summaries. Extensive ablations examine how time representation, marker frequency, tokenization, and duration-mixture design affect accuracy and computational cost. We release model weights and datasets to support further research on time-aware audio understanding, available at [https://huggingface.co/ai](https://huggingface.co/ai)sage/GigaChat3.1-Audio-10B-A1.8B.  \nIndex Terms: large audio language models, audio question answering, long-form speech processing  \n1. Introduction  \nLong recordings such as meetings, podcasts, lectures, and callcenter logs are increasingly consumed through interactive interfaces: users ask questions, request summaries, and navigate to specific evidence. In these settings, temporal grounding is not a cosmetic feature but a verifiability primitive: a system should answer not only what happened but also when it happened, enabling users to jump to the supporting audio segment.  \nRecent audio-conditioned LLMs enable instructionfollowing directly from speech and general audio, including both proprietary [1, 2] and open models [3, 4] . However, temporal grounding in long-form audio remains unreliable: models often produce plausible content while emitting non-parseable timestamps, overly coarse time references, or unsupported temporal claims. The core challenge is that time isnot naturally represented in standard audio token streams, and long recordings exacerbate the mismatch between the model’s internal representation and user-facing timestamped outputs.  \nThis work asks three practical questions: (i) What is the minimal and robust way to represent time in an audio LLM input/output interface? (ii) How should temporal supervision be generated at scale for long recordings, where manual annotation is prohibitively expensive? (iii) Do temporal models trained on a single duration regime generalize to other lengths? Our experiments show strong asymmetry: training only on short audio does not extrapolate to long recordings, while training only on long audio degrades short-audio performance (Figure 3) . Moreover, periodic temporal anchors are essential for longform grounding: removing them collapses long-recording accuracy (Table 1), although in practice even sparse anchors (e.g., once per minute) are sufficient (Table 3) .  \nFigure 1: Architecture of a time-aware audio LLM. Token streams are interleaved: text tokens, continuous audio tokens, timing tokens (either text-form hh:mm:ss or special ones).  \nWe present GigaChat Audio, a time-aware audioconditioned LLM that supports up to 120 minutes of input and produces time-anchored answers, fragment descriptions, and summaries with explicit timestamps. Our approach interleaves continuous audio tokens with periodic inter-timings that act as temporal anchors. To train temporal behavior at scale, we introduce a cascaded synthetic-data pipeline that generates supervision from timestamped transcripts using slicing to reduce frontloading bias and a global verifier to enforce consistency.  \nContributions.  \n• We release an open-we","cbCaimvTMNa93GbF","https://ap.wps.com/l/cbCaimvTMNa93GbF","pdf",635071,2,1,6,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does the GigaChat Audio model address?\",\"answer\":\"It tackles temporal grounding in long, audio-conditioned scenarios, where systems must answer not only what happened but also when it happened with explicit, verifiable timestamps.\"},{\"question\":\"How does the model represent time during audio processing?\",\"answer\":\"It interleaves continuous audio tokens with periodic inter-timing tokens that act as temporal anchors, supporting timestamped answers and time-anchored descriptions.\"},{\"question\":\"How is temporal supervision generated at scale?\",\"answer\":\"The approach uses a cascaded synthetic-data pipeline that builds supervision from timestamped transcripts, with transcript slicing to reduce bias and a global verifier to enforce consistency.\"}]",1784207013,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"gigachat-audio-time-aware-large-audio-language-model","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/gigachat-audio-time-aware-large-audio-language-model/85896/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the GigaChat Audio model address?","Question",{"text":75,"@type":76},"It tackles temporal grounding in long, audio-conditioned scenarios, where systems must answer not only what happened but also when it happened with explicit, verifiable timestamps.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the model represent time during audio processing?",{"text":80,"@type":76},"It interleaves continuous audio tokens with periodic inter-timing tokens that act as temporal anchors, supporting timestamped answers and time-anchored descriptions.",{"name":82,"@type":73,"acceptedAnswer":83},"How is temporal supervision generated at scale?",{"text":84,"@type":76},"The approach uses a cascaded synthetic-data pipeline that builds supervision from timestamped transcripts, with transcript slicing to reduce bias and a global verifier to enforce consistency.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]