[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86447-en":3,"doc-seo-86447-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86447,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","CLIR-Bench Benchmarking Multimodal Question Answering over Irregular Clinical Time Series","Clinical time series are central to patient monitoring, risk assessment, and clinical decision support, yet they are frequently sparse, irregularly sampled, and asynchronous across variables, complicating temporal grounding for clinical question answering. CLIR-Bench is introduced as an irregular clinical time-series QA benchmark built from de-identified ICU records via a principled four-stage pipeline. It includes 6,600 QA instances across 11 clinical variables, covering four capability dimensions and 11 tasks, with questions tied to explicit temporal evidence and answer-derivation rules.","CLIR-Bench: Benchmarking Multimodal Question Answering  \nover Irregular Clinical Time Series  \nFrank Nie∗ Shandong University China  \nEthan B. Liu∗†  \nShandong University China  \nYuan Zhu∗ Shandong University China  \n[zhu948982@gmail.com](zhu948982@gmail.com)  \nLoe Yan  \nShandong University China  \nWei Fan  \nUniversity of Auckland New Zealand  \nJindong Han† Shandong University China  \n[jindong.han@sdu.edu.cn](jindong.han@sdu.edu.cn)  \narXiv :2607 .09880v 1 [ cs .CL] 10 Jul 2026  \nAbstract  \nClinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irregularly sampled, and asynchronous, making it difficult for models to identify the temporal evidence required for clinical Question Answering (QA) . Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. CLIR-Bench contains 6,600 QA instances spanning 11 clinical variables, organized into four capability dimensions and 11 tasks. Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use. Experiments show that existing generalist models struggle to retrieve and reason over sparse clinical evidence, highlighting the need for stronger irregular time-series reasoning methods. Our code and data are available at [https://huggingface.co/datasets/winall/CLIR-Bench](https://huggingface.co/datasets/winall/CLIR-Bench).  \nKeywords  \nIrregular Clinical Time Series, Time Series Question Answering, Large Language Models  \nACM Reference Format:  \nFrank Nie, Ethan B. Liu, Yuan Zhu, Loe Yan, Wei Fan, and Jindong Han.  \n2018. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym ’XX) . ACM, New York, NY, USA, 14 pages. [https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \n∗ Equal Contribution.†Corresponding author.  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission [and/or a fee. Request permissions from permissions@acm.org](and/or a fee. Request permissions from permissions@acm.org).  \nConference acronym ’XX, Woodstock, NY  \n© 2018 Copyright held by the owner/author(s) . Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06  \n[https://doi.org/XXXXXXX.XXXXXXX](https://doi.org/XXXXXXX.XXXXXXX)  \n1 Introduction  \nClinical time series play a pivotal role in patient monitoring, risk assessment, and clinical decision support [22, 44] . In intensive care units (ICU), clinicians continuously interpret vital signs, laboratory measurements, and treatment events to understand how a patient’s condition changes over time [4, 19] . Unlike standard time series, however, clinical time series are often sparse, irregularly sampled, and asynchronous across variables [8]. Some measurements (e.g., heart rate) may be recorded every few minutes, while others, such as lactate, may appear only a few times over many hours. As a result, the information needed to answer a clinical question is often limited to a small number of observations scattered across alon","cbCaigqLsLKIhIrY","https://ap.wps.com/l/cbCaigqLsLKIhIrY","pdf",3432433,5,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"How does the benchmark evaluate whether a model uses the correct evidence?\",\"answer\":\"Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use.\"}]",1784211798,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"clir-bench-benchmarking-multimodal-question-answering-over-irregular-clinical-time-series","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/clir-bench-benchmarking-multimodal-question-answering-over-irregular-clinical-time-series/86447/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"How does the benchmark evaluate whether a model uses the correct evidence?","Question",{"text":76,"@type":77},"Each question is linked to explicit temporal evidence and task-specific answer derivation rules, enabling evaluation of both answer accuracy and evidence use.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":20,"slug":130},19,"General","general"]