[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123594-en":3,"doc-seo-123594-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123594,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","X-TIME - An in-memory engine for accelerating machine learning on tabular data with CAMs","Structured, or tabular, data is the dominant format in many data-science workflows, yet tree-based machine learning remains challenging to accelerate due to irregular memory access patterns and inference latency limits on CPUs and GPUs. X-TIME presents an analog-digital in-memory architecture that implements an increased-precision analog CAM combined with a programmable network-on-chip to run state-of-the-art tree models such as XGBoost and CatBoost. Single-chip evaluation on 16nm technology reports 119× lower latency with 9740× higher throughput versus a leading GPU, at 19W peak power.","X-TIME: An in-memory engine for accelerating machine learning on tabular data with CAMs  \nGiacomo Pedretti􀀃y , John Moon􀀃 , Pedro Bruel􀀃 , Sergey Serebryakov􀀃 , Ron M. Roth􀀃 , Luca Buonanno􀀃 , Tobias Ziegler􀀃 , Cong Xu􀀃 , Martin Foltin􀀃 , Paolo Faraboschi􀀃 , Jim Ignowski􀀃 , and Catherine E. Graves􀀃  \n􀀃 Artiﬁcial Intelligence Research Lab (AIRL), Hewlett Packard Labs, Milpitas (CA)  \ny giacomo.pedretti@hpe.com  \narXiv :2304 .0 1285v2 [ cs .LG] 5 Apr 2023  \nAbstract—Structured, or tabular, data is the most common format in data science. While deep learning models have proven formidable in learning from unstructured data such as images or speech, they are less accurate than simpler approaches when learning from tabular data. In contrast, modern tree-based Machine Learning (ML) models shine in extracting relevant information from structured data. An essential requirement in data science is to reduce model inference latency in cases where, for example, models are used in a closed loop with simulation to accelerate scientiﬁc discovery. However, the hardware acceleration community has mostly focused on deep neural networks and largely ignored other forms of machine learning. Previous work has described the use of an analog content addressable memory (CAM) component for efﬁciently mapping random forests. In this work, we focus on an overall analog-digital architecture implementing a novel increased precision analog CAM and a programmable network on chip allowing the inference of stateof-the-art tree-based ML models, such as XGBoost and CatBoost. Results evaluated in a single chip at 16nm technology show 119􀀂 lower latency at 9740􀀂 higher throughput compared with a stateof-the-art GPU, with a 19W peak power consumption.  \nI. INTRODUCTION  \nExtracting relevant information from structured (i.e., tabular) data is of utmost importance in practical data science. Tabular data consist of a matrix of samples organized in rows, each of which has the same set of features in its columns, and is commonly used for medical, ﬁnancial, scientiﬁc, and networking applications, enabling efﬁcient search operations on large-scale datasets. While deep learning has shown impressive improvements in accuracy for unstructured data, such as image recognition or machine translation, treebased machine learning (ML) models still outperform neural networks in classiﬁcation and regression tasks on tabular data [13], [21],[60] . This difference in model accuracy is due to robustness to uninformative features of tree-based models, bias towards a smooth solution of neural networks models, and in general the need for rotation invariant models where the data is rotation invariant itself, as in the case of tabular data [21] . With high model accuracy and ease of use, treebased ML models are greatly preferred by the data science community [7], with > 74% of data scientists preferring to use these models while \u003C 40% choose neural networks in a recent survey. However, tree-based models have garnered little attention in the architecture and custom accelerator research ﬁelds, particularly compared to deep learning. While small models are easily executed, performance issues for larger  \ntree-based ML models arise due to limited inference latency both on CPU and GPUs [26], [67] originating from irregular memory accesses and thread synchronization issues. Recent work attempted to address this limitation by improving GPU memory organization with compiler optimizations [4], [5] and model-tailored strategies for inference [67], with limited success. As large tree-based models are ﬁnding performance limits in existing hardware, larger models (>1M nodes) are demonstrating increasingly compelling uses in practical environments where high throughput and moderate latency are desired [1], [3] . For example, the winner of a recent IEEECIS Fraud detection competition used a tree-based ML model with 20M nodes [1] . Moreover, reducing inference latency is critical in scientiﬁc applications whe","cbCaitu1mTvAqjQs","https://ap.wps.com/l/cbCaitu1mTvAqjQs","pdf",1150088,1,13,"English","en",105,"# Introduction\n## Problem: latency and throughput limits for tree-based ML on tabular data\n## Prior work and gaps in hardware acceleration\n## Proposed approach: in-memory computing with analog CAMs\n## Target models and evaluation setup","[{\"question\":\"Why are tree-based machine learning models important for tabular data?\",\"answer\":\"Tree-based models extract relevant information from structured, tabular datasets and tend to outperform neural networks on many classification and regression tasks. They are also preferred in practice because of accuracy and usability on this data type.\"},{\"question\":\"What is the key bottleneck when accelerating large tree-based ML models on existing hardware?\",\"answer\":\"Inference latency and throughput limits arise from irregular memory accesses and thread synchronization issues, especially for larger trees with millions of nodes.\"},{\"question\":\"What does X-TIME contribute to accelerate tree-based ML inference?\",\"answer\":\"X-TIME introduces an analog-digital in-memory design using an increased-precision analog CAM plus a programmable network-on-chip. It enables efficient inference of models like XGBoost and CatBoost directly in memory, reducing data movement and lowering latency.\"}]","X-TIME - An in-memory engine for accelerating machine learning on tabular data with CAMs | PDF",1785817541,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"x-time-an-in-memory-engine-for-accelerating-machine-learning-on-tabular-data-with-cams","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/x-time-an-in-memory-engine-for-accelerating-machine-learning-on-tabular-data-with-cams/123594/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are tree-based machine learning models important for tabular data?","Question",{"text":75,"@type":76},"Tree-based models extract relevant information from structured, tabular datasets and tend to outperform neural networks on many classification and regression tasks. They are also preferred in practice because of accuracy and usability on this data type.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key bottleneck when accelerating large tree-based ML models on existing hardware?",{"text":80,"@type":76},"Inference latency and throughput limits arise from irregular memory accesses and thread synchronization issues, especially for larger trees with millions of nodes.",{"name":82,"@type":73,"acceptedAnswer":83},"What does X-TIME contribute to accelerate tree-based ML inference?",{"text":84,"@type":76},"X-TIME introduces an analog-digital in-memory design using an increased-precision analog CAM plus a programmable network-on-chip. It enables efficient inference of models like XGBoost and CatBoost directly in memory, reducing data movement and lowering latency.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]