[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86451-en":3,"doc-seo-86451-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86451,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","A Foundation Model for Multimodal Event Sequences in Financial Applications","Predictive modeling is a core component of modern financial services, where many tasks are traditionally handled by separate models trained on manually engineered tabular features. This task-specific design limits reuse and underutilizes heterogeneous sources such as transaction histories and digital interaction signals. The approach pretrains a foundation transformer on multimodal user event sequences, unifying events into a chronological timeline and using next-event prediction for self-supervised learning. Learned representations are combined with engineered features to drive multiple downstream tasks, outperforming task-specific baselines while reducing development overhead.","A Foundation Model for Multimodal Event Sequences in Financial  \nApplications  \nNikita Rusakov∗ Sber  \nMoscow, Russian Federation [nanrusakov@gmail.com](nanrusakov@gmail.com)  \nGleb Zaripov∗  \nSber  \nMoscow, Russian Federation [gleb.zar.030@gmail.com](gleb.zar.030@gmail.com)  \nVladislav Meshkov∗ Sber  \nMoscow, Russian Federation [vladmeshkov160@gmail.com](vladmeshkov160@gmail.com)  \nAlexander Uglov∗  \nSber  \nMoscow, Russian Federation [skilletthebest25@gmail.com](skilletthebest25@gmail.com)  \nKonstantin Zorin∗ Sber Moscow, Russian Federation [kostya.zorin.2003.ri@gmail.com](kostya.zorin.2003.ri@gmail.com)  \nAlexey Vasilev  \nSber AI Lab Moscow, Russian Federation [alexxl.vasilev@yandex.ru](alexxl.vasilev@yandex.ru)  \nAnton Klenitskiy  \nSber AI Lab Moscow, Russian Federation  \n[antklen@gmail.com](antklen@gmail.com)  \narXiv :2607 .09955v 1 [ cs .LG] 10 Jul 2026  \nAbstract  \nPredictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separate models trained on manually engineered tabular features. This task-specific approach limits reuse and makes it difficult to fully exploit heterogeneous data sources such as transaction histories and digital interaction signals. In this paper, we present an approach based on pretraining a foundation transformer model on multimodal sequences of user events. Events from multiple data sources are unified into a single chronological sequence, enabling early fusion of heterogeneous modalities and learning of generalpurpose representations via a next-event prediction objective. These representations are combined with existing engineered user features, on top of which lightweight neural models are trained for multiple downstream tasks. The proposed system outperforms traditional task-specific models while reducing development overhead. The approach was deployed in production at one of the biggest banks in Eastern Europe, resulting in measurable improvements in business metrics.  \nCCS Concepts  \n• Information systems → Data mining; • Computing methodologies → Neural networks; Learning latent representations.  \nKeywords  \nfoundation models; multimodal event sequences; self-supervised learning; transactional data; financial applications  \n∗ Authors contributed equally to the paper  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License. KDD’26, Jeju Island, Republic of Korea  \n© 2026 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-2259-2/2026/08  \n[https://doi.org/10.1145/3770855.3818311](https://doi.org/10.1145/3770855.3818311)  \nACM Reference Format:  \nNikita Rusakov, Vladislav Meshkov, Konstantin Zorin, Gleb Zaripov, Alexander Uglov, Alexey Vasilev, and Anton Klenitskiy. 2026. A Foundation Model for Multimodal Event Sequences in Financial Applications. In Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 (KDD’26), August 09–13, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 9 pages. [https://doi.org/10.1145/3770855.3818311](https://doi.org/10.1145/3770855.3818311)  \n1 Introduction  \nMachine learning is a core component of modern financial services, supporting a wide range of predictive tasks such as risk assessment, fraud detection, product recommendations, and customer analytics [1, 3, 11, 30, 35] . In large financial institutions, these tasks are typically addressed by building separate models for each use case, trained on manually engineered features derived from user transaction histories and digital interaction data. While this task-specific approach has demonstrated strong performance in individual applications, it leads to significant duplication of effort, limited reuse of learned representations, and increasing system complexity asthe number of predictive tasks grows.  \nA key limitation of this paradigm is its inability to fully leverage the scale and richness of available data [3, 35] . In practice, user behavior is record","cbCaiuNkn5fuoEL2","https://ap.wps.com/l/cbCaiuNkn5fuoEL2","pdf",698338,7,1,9,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why do traditional task-specific models limit financial analytics in practice?\",\"answer\":\"They rely on manually engineered tabular features and separate models per use case, which reduces representation reuse and makes it difficult to exploit heterogeneous event data fully.\"},{\"question\":\"How does the proposed approach pretrain the foundation model?\",\"answer\":\"It unifies events from multiple data sources into a single chronological multimodal sequence and trains a transformer using a self-supervised next-event prediction objective.\"},{\"question\":\"How are the pretrained representations used for downstream tasks?\",\"answer\":\"The model embeddings serve as a shared backbone that is combined with existing engineered user features, on top of which lightweight task-specific neural models are trained.\"}]",1784211819,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-foundation-model-for-multimodal-event-sequences-in-financial-applications","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-foundation-model-for-multimodal-event-sequences-in-financial-applications/86451/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do traditional task-specific models limit financial analytics in practice?","Question",{"text":76,"@type":77},"They rely on manually engineered tabular features and separate models per use case, which reduces representation reuse and makes it difficult to exploit heterogeneous event data fully.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed approach pretrain the foundation model?",{"text":81,"@type":77},"It unifies events from multiple data sources into a single chronological multimodal sequence and trains a transformer using a self-supervised next-event prediction objective.",{"name":83,"@type":74,"acceptedAnswer":84},"How are the pretrained representations used for downstream tasks?",{"text":85,"@type":77},"The model embeddings serve as a shared backbone that is combined with existing engineered user features, on top of which lightweight task-specific neural models are trained.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]