[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85250-en":3,"doc-seo-85250-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85250,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos","Assembly action understanding is essential for effective human-robot collaborative assembly, yet it is difficult because of subtle motion, fine-grained hand–object interactions, and actions that vary in motion or partners. The work adapts vision-language models using Compositional Context Fine-Tuning (CCFT), decomposing actions into Verb/Object/Tool elements and fine-tuning with templated VQA pairs for near-deterministic outputs. A Layer-Partitioned Alternating Training (LP-AT) strategy assigns disjoint layer groups to element-specific LoRA adapters to reduce cross-task interference under limited data. HA-ViD-VQA and IKEA-ASM-VQA datasets support experiments showing consistent gains over action recognition baselines and interpretable element-level predictions.","Compositional Context Fine-Tuning Vision-Language Model for Complex Assembly Action Understanding from Videos  \nHao Zheng∗ , 1 ,4 , Jinyi Huang4 , Tiantian Zheng 1 ,2 , Xun Xu4 , Tuka Alhanai 1 ,3  \narXiv :2607 . 10797v1 [ cs .CV] 12 Jul 2026  \nAbstract—Assembly action understanding is a key enabler for effective human-robot collaborative assembly, yet it remains challenging due to subtle motions and fine-grained hand–object interactions. We adapt vision-language models (VLMs) to this challenging domain with Compositional Context Fine-Tuning (CCFT), a method that decomposes assembly actions into semantic elements (Verb, Object, Tool) and fine-tunes VLMs to recognize each action element using templated questionanswering pairs. This approach ensures near-deterministic outputs. To enable efficient and effective multi-task learning under limited data, a Layer-Partitioned Alternating Training (LP-AT) method is presented, which assigns distinct model layers to recognize specific action elements through elementspecific low-rank adapters. LP-AT alternates weight updates across element-specific adapters, reducing cross-task interference while enabling per-adapter hyperparameter optimization. Furthermore, we create HA-ViD-VQA and IKEA-ASM-VQA datasets from existing assembly video datasets. Extensive experiments on these datasets demonstrate that our method consistently outperforms strong action recognition baselines while providing interpretable element-level predictions that can support diverse downstream applications. Code and dataset are released at [https://github.com/x-labs-xyz/CCFT](https://github.com/x-labs-xyz/CCFT).  \nI. INTRODUCTION  \nBuilding on the strong textual reasoning of large language models (LLMs), vision–language models (VLMs) have advanced multimodal understanding by integrating visual and linguistic cues [1],[2] . Multimodal perception enables VLMs to tackle a broader range of real-world challenges, positioning them as promising foundations for human-robot collaboration (HRC), which demands highly precise environmental understanding and physical interaction.  \nAssembly action understanding represents a critical component of collaborative robotics, as robots must comprehend complex assembly actions to effectively assist humans and acquire manipulation skills through demonstration [3],[4] . However, assembly action understanding poses unique challenges for robotic perception: subtle and intricate actions, distinct actions with similar motions, identical actions with distinct motions, and nuanced hand-object interactions. While VLMs offer considerable potential for advancing complex assembly action understanding, this application domain  \n∗ Hao Zheng is corresponding author: [h.zheng@nyu.edu](h.zheng@nyu.edu);  \n1Department of Computer Engineering, New York University Abu Dhabi, UAE; 2Center for Quantum and Topological Systems, New York University Abu Dhabi, UAE; 3 Center for AI and Robotics, NYUAD, UAE; 4 Department of Mechanical and Mechatronics Engineering, The University of Auckland, New Zealand. H.Z.: [hzhe951@aucklanduni.ac.nz](hzhe951@aucklanduni.ac.nz);  \nJ.H: [jhua658@aucklanduni.ac.nz](jhua658@aucklanduni.ac.nz); X.X: [x.xu@auckland.ac.nz](x.xu@auckland.ac.nz. T.A)[. T.A](x.xu@auckland.ac.nz. T.A) acknowledges support by CAIRand CQTS funded by Tamkeen NYUAD Research Institute Award CG010 and CG008, respectively.  \nremains largely unexplored.  \nMost existing VLM research targets general scene understanding, supporting only basic video summarization or simple question-answering (QA) tasks [5], and lacking specialized adaptions for recognizing complex assembly actions. Moreover, HRC applications demand deterministic outputs, which conflicts with the generative nature of VLM outputs.  \nThis paper proposes two complementary technical ingredients to address these challenges. First, a Compositional Context Fine-Tuning (CCFT) method is proposed to decompose assembly actions into semantic elements—Verb, Object, and T","cbCaifXrhov67Yx6","https://ap.wps.com/l/cbCaifXrhov67Yx6","pdf",774217,3,1,"English","en",105,"# Introduction\n## Problem and motivation\n## Proposed approach: CCFT\n## Proposed training approach: LP-AT\n## Dataset reformulation and evaluation","[{\"question\":\"What is CCFT in this work?\",\"answer\":\"CCFT decomposes assembly actions into semantic elements (Verb, Object, Tool) and fine-tunes a vision-language model to recognize each element using templated VQA question-answer pairs for near-deterministic predictions.\"},{\"question\":\"How does LP-AT improve multi-task learning with limited data?\",\"answer\":\"LP-AT assigns disjoint model layer groups to different action elements and trains element-specific LoRA adapters by alternating updates, limiting gradient interference across tasks while keeping parameter-efficient training.\"},{\"question\":\"What datasets are introduced or created, and why?\",\"answer\":\"The work reformulates existing assembly video datasets (HA-ViD and IKEA-ASM) into compositional VQA datasets (HA-ViD-VQA and IKEA-ASM-VQA) with templated QA pairs and element-level annotations to validate compositional action understanding.\"}]",1784202072,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"compositional-context-fine-tuning-vision-language-model-for-complex-assembly-action-understanding-from-videos","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/compositional-context-fine-tuning-vision-language-model-for-complex-assembly-action-understanding-from-videos/85250/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What is CCFT in this work?","Question",{"text":74,"@type":75},"CCFT decomposes assembly actions into semantic elements (Verb, Object, Tool) and fine-tunes a vision-language model to recognize each element using templated VQA question-answer pairs for near-deterministic predictions.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does LP-AT improve multi-task learning with limited data?",{"text":79,"@type":75},"LP-AT assigns disjoint model layer groups to different action elements and trains element-specific LoRA adapters by alternating updates, limiting gradient interference across tasks while keeping parameter-efficient training.",{"name":81,"@type":72,"acceptedAnswer":82},"What datasets are introduced or created, and why?",{"text":83,"@type":75},"The work reformulates existing assembly video datasets (HA-ViD and IKEA-ASM) into compositional VQA datasets (HA-ViD-VQA and IKEA-ASM-VQA) with templated QA pairs and element-level annotations to validate compositional action understanding.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]