[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83955-en":3,"doc-seo-83955-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83955,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation","Vision-Language-Action (VLA) models show promise for generalist manipulation, yet long-horizon tasks fail because Markovian decision-making relies only on current observations. Cortex introduces a bidirectionally aligned embodied agent framework that bridges high-level planning semantics and low-level execution kinematics via a customized planning interface. It standardizes manipulation into 32 canonical skill primitives, injects tractability principles into data generation, and uses event-balanced sampling for fine-tuning, including harness engineering for inference. Results show improved success on Liberolong and RoboTwin and zero-shot real-world long-horizon completion by combining with a fine-tuned VLA.","arXiv :2607 .05377v 1 [ cs .RO] 6 Jul 2026  \nCortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation  \nJiaqi Peng∗ 1 ,2 Xiqian Yu∗ 2 Delin Feng∗ 2 Yuqiang Yang2 Wenzhe Cai2 Jing Xiong2 ,3 Ganlin Yang2 ,4 Jinliang Zheng 1 ,2 Jiafei Cao2 Xueyuan Wei2 Jiangmiao Pang2 Yuan Shen† 1 Tai Wang† 2  \n1Tsinghua University 2 Shanghai AI Laboratory 3Peking University 4USTC  \n[https://steinate.github.io/cortex.github.io](https://steinate.github.io/cortex.github.io)  \n∗Equal contribution. †Corresponding author.  \nAbstract: While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature—relying solely on current observations. Hierarchical dual-system methods address this but suffer from a gap between high-level planning semantics and low-level execution kinematics. We introduce Cortex, a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from high-level VLM to low-level VLA. Specifically, we standardize manipulation subtasks into 32 canonical skill primitives and inject tractability principles, such as representative object attributes and improved trajectory reachability, into the data generation pipeline.  \nThis enables automatic annotation of over 4k hours of open-source video data and generation of 30 hours of simulation data. We further devise an event-balanced sampling strategy to construct training data for fine-tuning the framework to better handle planning ambiguity during subtask transitions, enhanced by carefully designed harness engineering from task contexts to skill constraints during inference. Both open-loop VLM and closed-loop system evaluations demonstrate Cortex’s efficacy, e.g., it outperforms monolithic baselines by 3.1% on Liberolong and 4.1% on RoboTwin. Notably, Cortex’s generalist VLM enables zero-shot completion of unseen real-world long-horizon tasks, such as multi-stage chemistry experiments, by simply combining with a fine-tuned VLA—a capability infeasible through VLA fine-tuning alone.  \nKeywords: Long-horizon Manipulation, Vision-Language-Action Model  \nGlobal Instruction: Wash the beaker  \n􀁘 􀁙 􀁚 􀁛 􀁜 􀁝 􀁞 􀁟 􀁠  \n◼ Pick up the beaker ◼ Place the beaker on the platform ◼ Pick up the bottle ◼ Unscrew the cap ◼ Pour the water into the beaker  \n◼ Place the bottle on the table ◼ Pick up the beaker ◼ Pour the water into the kettle ◼ Place the beaker on the table  \nFigure 1: Compared to previous works, Cortex is an embodied agent framework aligning the highlevel VLM and low-level VLA on the subtask executability and tractability, enabling synergistic planning and execution for long-horizon manipulation.  \n1 Introduction  \nVision-Language-Action (VLA) models have fundamentally reshaped embodied AI by directly mapping multi-modal inputs to continuous motor control, achieving remarkable zero-shot generalization in short-horizon tasks [1, 2, 3, 4] . However, as tasks scale in temporal complexity, monolithic VLAs face a critical bottleneck: Markovian short-sightedness. Operating purely reactively without continuous progress verification or spatial-temporal memory, these models struggle to differentiate actual task progress from instantaneous visual observations [5, 6] . Consequently, when executing monolithic long-horizon instructions, they frequently lose track of intermediate states and blindly repeat actions, inevitably leading to compounding execution errors.  \nRecent efforts to mitigate this via visual frame buffering [7, 8] are constrained by limited context windows and struggle to form the semantic memory required for logical planning. Alternatively, hierarchical dual-system paradigms decouple cognitive planning from reactive execution. However, they fundamentally suffer from a lack of bidirectional alignment, i.e., the high-level VLM should consider the capabilities of low-level VLA, while the VLA","cbCaihQxWVn17mfv","https://ap.wps.com/l/cbCaihQxWVn17mfv","pdf",40953927,5,1,33,"English","en",105,"# Introduction\n## Background: Limitations of Markovian monolithic VLA\n## Hierarchical dual-system gap and misalignment\n## Cortex: bidirectionally aligned framework","[{\"question\":\"Why do monolithic VLA models struggle with long-horizon manipulation?\",\"answer\":\"They are Markovian and operate reactively on current observations without reliable progress verification or semantic memory. This can cause repeated or incorrect intermediate actions, leading to compounding execution errors.\"},{\"question\":\"What core mechanism does Cortex use to align planning and execution?\",\"answer\":\"Cortex uses a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from a high-level VLM to a low-level VLA.\"},{\"question\":\"How does Cortex improve training data quality and planning-transition robustness?\",\"answer\":\"It standardizes manipulation subtasks into 32 canonical skill primitives for automatic annotation, injects physical tractability principles into data generation (including improved trajectory reachability), and uses event-balanced sampling to build subtask execution/transition training samples for fine-tuning.\"}]",1784191652,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"cortex-a-bidirectionally-aligned-embodied-agent-framework-for-long-horizon-manipulation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/cortex-a-bidirectionally-aligned-embodied-agent-framework-for-long-horizon-manipulation/83955/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do monolithic VLA models struggle with long-horizon manipulation?","Question",{"text":76,"@type":77},"They are Markovian and operate reactively on current observations without reliable progress verification or semantic memory. This can cause repeated or incorrect intermediate actions, leading to compounding execution errors.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What core mechanism does Cortex use to align planning and execution?",{"text":81,"@type":77},"Cortex uses a bidirectionally aligned embodied agent framework with a customized planning interface that conveys executable and tractable subtask plans from a high-level VLM to a low-level VLA.",{"name":83,"@type":74,"acceptedAnswer":84},"How does Cortex improve training data quality and planning-transition robustness?",{"text":85,"@type":77},"It standardizes manipulation subtasks into 32 canonical skill primitives for automatic annotation, injects physical tractability principles into data generation (including improved trajectory reachability), and uses event-balanced sampling to build subtask execution/transition training samples for fine-tuning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]