[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151701-en":3,"doc-seo-151701-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151701,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","OLMOTRACE - Tracing Language Model Outputs","OLMOTRACE is a real-time system that traces language model outputs back to multi-trillion-token training data. It identifies and displays verbatim matches between spans in generated text and corresponding documents in the training corpora, using an extended infini-gram indexing approach for fast exact-match retrieval. The system returns tracing results within seconds, enabling interactive exploration of behaviors such as fact checking, hallucination, and creativity through the lens of training evidence. OLMOTRACE is open-source and publicly available.","OLMOTRACE: Tracing Language Model Outputs  \nBack to Trillions of Training Tokens  \nJiacheng Liuαω Taylor Blantonα Yanai Elazarαω Sewon Minαβ YenSung Chenα Arnavi Chheda-Kotharyαω Huy Tranα Byron Bischoffα Eric Marshα Michael Schmitzα Cassidy Trierα Aaron Sarnatα Jenna Jamesα Jon Borchardtα Bailey Kuehlα Evie Chengα Karen Farleyα Sruthi Sreeramα Taira Andersonα David Albrightα Carissa Schoenickα Luca Soldainiα Dirk Groeneveldα  \nRock Yuren Pangω  \nPang Wei Kohαω Noah A. Smithαω Sophie Lebrechtα Yejin Choiσ Hannaneh Hajishirziαω Ali Farhadiαω Jesse DodgeααAllen Institute for AI ω University of Washington β UC Berkeley σ Stanford University  \nAbstract  \nWe present OLMOTRACE, the first system that traces the outputs of language models back to their full, multi-trillion-token training data in real time. OLMOTRACE finds and shows verbatim matches between segments of language model output and documents in the training text corpora. Powered by an extended version of infini-gram (Liu et al., 2024), our system returns tracing results within a few seconds. OLMOTRACE can help users understand the behavior of language models through the lens of their training data. We showcase how it can be used to explore fact checking, hallucination, and the creativity of language models.  \nOLMOTRACE is publicly available and fully open-source.  \n1 Introduction  \nTracing the outputs of language models (LMs) back to their training data is an important problem. As LMs gain adoption in higher-stakes scenarios, it is critical to understand why they generate certain responses. However, these modern LMs are trained on massive text corpora with trillions of tokens, which are often proprietary. Fully open LMs (e.g., OLMo; OLMo et al. 2024) enable access to the training data, but existing behavior tracing methods (Koh and Liang, 2017 ; Khalifa et al., 2024 ; Huang et al., 2024) have not been scaled to work within this multi-trillion-token setting due to their heavy computational needs.  \nFigure 1: OLMOTRACE on Ai2 Playground. Left: On a response generated by OLMo, OLMOTRACE highlights text spans found verbatim in the model’s training data and shows their source documents. Brighter highlights indicate spans from more relevant training documents, while darker highlights denote less relevant ones. Right:  \nWhen user clicks the “View Document” button, the document is shown with extended context. Try OLMOTRACE at [https://playground.allenai.org](https://playground.allenai.org).  \n178  \nProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations) , pages 178–188  \nJuly 27-August 1, 2025 ©2025 Association for Computational Linguistics  \nIn this paper, we introduce OLMOTRACE, a system that traces LM outputs verbatim back to its full training data and displays the tracing results to LM users in real time. Given an LM response toa user prompt, OLMOTRACE retrieves documents from the model’s training data that contain exact matches with pieces of the LM response that are long, unique, and relevant to the whole response; see Figure 1 for an example.  \nThe key idea that makes OLMOTRACE fast is that exact matches can be quickly located in a large text corpus if we pre-sort all of its suffixes lexicographically. We use infini-gram (Liu et al., 2024) to index the training data and develop a novel parallel algorithm to speed up the computation of matching spans (§3) . In our production system, OLMOTRACE completes tracing for each LM response (avg. ∼450 tokens) within 4.5 seconds on average. The purpose of OLMOTRACE is to give users a tool to explore where LMs may have learned to generate certain word sequences, focusing on verbatim matching as the most direct connection between LM outputs and the training data. OLMOTRACE offers an interactive experience, so that users can explore which training documents contain a specific span in the LM response, or inspect a particular document and locate its matching spansin the LM re","cbCaikBliE6cfNKW","https://ap.wps.com/l/cbCaikBliE6cfNKW","pdf",4960547,1,11,"English","en",105,"# Introduction\n## Problem and motivation\n## Key idea and system performance\n# System Description\n## Features and workflow","[{\"question\":\"What does OLMOTRACE trace and how are matches found?\",\"answer\":\"OLMOTRACE traces language model outputs back to training data by locating verbatim matches between spans in a response and segments in the training corpora. It uses an extended infini-gram approach for fast exact-match retrieval.\"},{\"question\":\"How fast are tracing results returned?\",\"answer\":\"In the production system, OLMOTRACE completes tracing for each language model response (about 450 tokens on average) within about 4.5 seconds on average.\"},{\"question\":\"What can users do with OLMOTRACE after enabling it?\",\"answer\":\"Users can explore which training documents contain highlighted spans in the response, or inspect a particular document to locate where its matching spans appear in the language model output. The paper also demonstrates use cases including fact checking, hallucination/behavior analysis, and tracing creative and mathematical capabilities.\"}]","OLMOTRACE - Tracing Language Model Outputs | PDF",1787847569,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"olmotrace-tracing-language-model-outputs","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/olmotrace-tracing-language-model-outputs/151701/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does OLMOTRACE trace and how are matches found?","Question",{"text":76,"@type":77},"OLMOTRACE traces language model outputs back to training data by locating verbatim matches between spans in a response and segments in the training corpora. It uses an extended infini-gram approach for fast exact-match retrieval.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How fast are tracing results returned?",{"text":81,"@type":77},"In the production system, OLMOTRACE completes tracing for each language model response (about 450 tokens on average) within about 4.5 seconds on average.",{"name":83,"@type":74,"acceptedAnswer":84},"What can users do with OLMOTRACE after enabling it?",{"text":85,"@type":77},"Users can explore which training documents contain highlighted spans in the response, or inspect a particular document to locate where its matching spans appear in the language model output. The paper also demonstrates use cases including fact checking, hallucination/behavior analysis, and tracing creative and mathematical capabilities.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]