[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84516-en":3,"doc-seo-84516-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84516,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","ClothTransformer Unified Latent-Space Transformers for Scalable Cloth Simulation","Unified and scalable Transformers have proven effective for modeling diverse computer-graphics and video phenomena. ClothTransformer extends this capability to cloth simulation by reformulating cloth dynamics as autoregressive sequence modeling in a learned latent space. The framework delivers a unified model for multiple scenarios, a scalable latent representation that decouples computation from mesh resolution, and a large penetration-free dataset. It further incorporates differentiable continuous collision detection to reduce penetration artifacts, achieving substantially lower prediction error across settings.","ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation  \nYu Zhang 1 Yidi Shao2 Wenqi Ouyang 1 Yushi Lan3 Zhexin Liang 1 Chengrui Wu4 Xudong Xu5 Xingang Pan 1  \n1 S-Lab, Nanyang Technological University, Singapore 2Feeling AI  \n3University of Oxford 4Nanyang Technological University 5 Shanghai AI Laboratory  \narXiv :2605 .27852v4 [ cs .GR] 13 Jul 2026  \nFigure 1: ClothTransformer generalizes to unseen test cases across three diverse scenarios. Left two: Diverse Object Collision—cloth falling onto unseen rigid objects (sword, character) . Middle two: Human Garment—unseen body, garment, and animation combinations (front-flip, dancing) . Right two: Robotic Manipulation—unseen cloth meshes grasped and lifted by a robotic gripper.  \nAbstract  \nUnified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graphics, such as 3D visual effects, rendering processes, and motion in videos. In this work, we take a step further by investigating whether modern Transformer techniques can tackle the challenging task of cloth simulation. To this end, we present ClothTransformer, a framework that reformulates cloth simulation as autoregressive sequence modeling in a learned latent space. Existing neural cloth simulators are largely specialized to single scenarios, intrinsically coupled to the mesh discretization, and lack robust collision handling. Our approach addresses these limitations through three contributions: (1) a unified Transformer architecture that handles diverse scenarios—body-driven garments, robotic manipulation, and freefall collisions—under a single model and achieves approximately 4–9× lower error than prior state-of-the-art methods across all scenarios; (2) a scalable latent-space formulation that compresses arbitrary-resolution meshes into a fixed-size set of latent tokens, making temporal dynamics computation independent of mesh resolution; and (3) a diverse-scenario high-fidelity penetration-free dataset of ∼493.4k frames spanning all three settings, which enables a differentiable Continuous Collision Detection (CCD) module to suppress penetration artifacts. Project Page:  \n[https://yucrazing.github.io/clothtransformer/](https://yucrazing.github.io/clothtransformer/)  \nPreprint.  \n1 Introduction  \nRealistic cloth simulation is essential for a wide range of applications. In film and visual effects, convincing fabric motion brings digital characters to life; in gaming and virtual reality, interactive garments are key to immersion; and in embodied AI, the recent rapid development further intensifies the demand for efficient and physically plausible simulation. Despite decades of progress, however, simultaneously achieving high fidelity and real-time performance remains challenging. Physically Based Simulation (PBS) methods [1], including advanced variational contact solvers such as IPC [17], can produce highly accurate results; yet even with modern GPU acceleration [15], high-resolution cloth can still take tens of seconds per frame—far beyond real-time budgets.  \nLearning-based neural simulators offer a promising alternative. Most recent progress is driven by Graph Neural Networks (GNNs) [28, 9], which predict vertex dynamics via message passing on mesh edges. Still, existing approaches face three fundamental limitations:  \nLack of Generalization. Existing learning-based cloth simulators are largely specialized to a single setting—typically human-garment dressing on an animated body—or require training a separate model for each scenario. In both cases, they lack a single unified architecture and model capable of handling diverse scenarios such as robotic manipulation or free-fall collisions, which hinders their applicability to broader simulation tasks.  \nThe Resolution Bottleneck. GNN simulators are tightly coupled to mesh discretization: inference cost grows with vertex/edge count. This creates a direct conflict bet","cbCaitXk2gmkVfb9","https://ap.wps.com/l/cbCaitXk2gmkVfb9","pdf",39038246,1,20,"English","en",105,"# Introduction\n## Motivation and Challenges\n## Core Idea: Unified Transformer in Latent Space\n## Contributions and Technical Components","[{\"question\":\"What is ClothTransformer’s core modeling approach for cloth simulation?\",\"answer\":\"ClothTransformer reformulates cloth simulation as autoregressive sequence modeling in a learned latent space, evolving cloth dynamics using Transformer techniques rather than direct mesh-based prediction.\"},{\"question\":\"How does ClothTransformer address the resolution bottleneck?\",\"answer\":\"It compresses arbitrary-resolution cloth meshes into a fixed-size set of latent tokens and performs temporal computation in latent space, making runtime effectively independent of mesh resolution.\"},{\"question\":\"How does ClothTransformer reduce penetration artifacts during training and inference?\",\"answer\":\"It uses a diverse-scenario penetration-free dataset to enable a differentiable continuous collision detection (CCD) loss during training and CCD post-processing at inference, suppressing tunneling from discrete checks.\"}]",1784196253,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"clothtransformer-unified-latent-space-transformers-for-scalable-cloth-simulation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/clothtransformer-unified-latent-space-transformers-for-scalable-cloth-simulation/84516/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is ClothTransformer’s core modeling approach for cloth simulation?","Question",{"text":75,"@type":76},"ClothTransformer reformulates cloth simulation as autoregressive sequence modeling in a learned latent space, evolving cloth dynamics using Transformer techniques rather than direct mesh-based prediction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ClothTransformer address the resolution bottleneck?",{"text":80,"@type":76},"It compresses arbitrary-resolution cloth meshes into a fixed-size set of latent tokens and performs temporal computation in latent space, making runtime effectively independent of mesh resolution.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ClothTransformer reduce penetration artifacts during training and inference?",{"text":84,"@type":76},"It uses a diverse-scenario penetration-free dataset to enable a differentiable continuous collision detection (CCD) loss during training and CCD post-processing at inference, suppressing tunneling from discrete checks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":28,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]