[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84303-en":3,"doc-seo-84303-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84303,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","APIVOT Adaptive Planning with Interleaved Vision-Language Thoughts","Long-horizon robotic planning requires combining semantic task understanding with geometric feasibility to decompose goals, select relevant objects, and order actions under spatial constraints such as limited free space and object collisions. APIVOT is a VLM-based planner that adaptively interleaves language-based reasoning with visual thoughts that represent imagined future states for internal geometric verification. On long-horizon kitchen tasks, APIVOT improves performance over general VLMs and prior planning methods, with the largest gains in spatially constrained settings, while increasing planning success and reasoning efficiency.","arXiv :2607 .08024v 1 [ cs .CV] 9 Jul 2026  \nAPIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts  \nEmily Jin Joy Hsu Yiqing Xu Weiyu Liu† Nick Haber† Jiajun Wu†  \nStanford University  \nAbstract  \nLong-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions.  \nIn this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to leverage language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency. *  \n1 Introduction  \nLong-horizon robot planning requires flexibly interleaving semantic and geometric reasoning. Consider the task: “store the leftovers in the fridge.” To successfully achieve this goal, a robot must reason semantically to identify which items need to be stored, select appropriate containers, and determine a sequence of actions that satisfies prerequisites (e.g., the fridge must be open before placing anything inside) . At the same time, successful execution depends on geometric constraints, such as whether the leftovers fit inside the selected containers, how the containers should be arranged inside the fridge, and whether any existing objects must be moved aside to create enough free space. These two modes of reasoning are deeply intertwined. A symbolically valid plan may still fail if the containers collide, while geometric decisions made early on can affect which actions remain feasible downstream. Existing LLM-and VLM-based planners can reliably produce semantically plausible action sequences, but often struggle when success depends on geometric feasibility [1–8] . Prior work addresses this by coupling the model with external motion planners, feasibility checkers, or learned dynamics models [9– 13] . However, these systems typically incorporate geometric feedback for replanning, which does not shape the planner’s internal reasoning. Instead, we argue that an effective planner should interleave semantic and geometric reasoning itself, using language and vision as complementary modalities. Language is effective for semantic reasoning, such as task decomposition and action selection [14], but it cannot express the geometric structure needed to plan over resulting scene configurations in a compact, precise way. By contrast, visual representations effectively encode spatial layout, object shapes, and remaining free space, making them well-suited for geometric reasoning [15, 16] . Thus, language and vision are complementary modes of reasoning that should be used adaptively based on a task’s demands. While recent vision-language-action (VLA) models begin to incorporate visual representations for intermediate reasoning, they apply them uniformly, and a challenge remains in learning when to reason in vision and language [14–19] .  \n†Equal advising.  \n* Project page: [https://emilyzjin.github.io/projects/apivot.html](https://emilyzjin.github.io/projects/apivot.html)  \nPreprint.  \nInputs Adaptive Planning with Interleaved Vision-Language Thoughts  \nObjects:  \nfridge0, bowl1, ...  \n| \u003Cbr>VLM |  | Open fridge0 . bowl1 is already inside. To ensure that none of the objects collide, place bowl2 in front of bowl1 and on the left side of the fridge. Place bowl3 ... |  | Plan: Open(fridge0) Pick(bowl2), Place(bowl2,[3,8]),\u003Cbr>Pick(bowl3), Plac","cbCaips6KgSajnkz","https://ap.wps.com/l/cbCaips6KgSajnkz","pdf",5878401,5,1,37,"English","en",105,"# Abstract\n# Introduction\n# Method","[{\"question\":\"What problem does APIVOT target in long-horizon robot planning?\",\"answer\":\"APIVOT targets the need to jointly reason about semantic task structure and geometric feasibility so that action sequences satisfy spatial constraints like limited free space and potential object collisions.\"},{\"question\":\"How does APIVOT use language and visual thoughts during planning?\",\"answer\":\"APIVOT uses language for semantic reasoning such as subgoal decomposition and action ordering, while visual thoughts represent imagined future states to internally verify geometric feasibility before execution.\"},{\"question\":\"Why does APIVOT outperform general-purpose VLMs in the reported experiments?\",\"answer\":\"It achieves larger gains in spatially constrained settings by learning adaptive modality selection, choosing whether to reason in vision or language at each step to improve both planning success and reasoning efficiency.\"}]",1784194689,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"apivot-adaptive-planning-with-interleaved-vision-language-thoughts","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/apivot-adaptive-planning-with-interleaved-vision-language-thoughts/84303/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does APIVOT target in long-horizon robot planning?","Question",{"text":76,"@type":77},"APIVOT targets the need to jointly reason about semantic task structure and geometric feasibility so that action sequences satisfy spatial constraints like limited free space and potential object collisions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does APIVOT use language and visual thoughts during planning?",{"text":81,"@type":77},"APIVOT uses language for semantic reasoning such as subgoal decomposition and action ordering, while visual thoughts represent imagined future states to internally verify geometric feasibility before execution.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does APIVOT outperform general-purpose VLMs in the reported experiments?",{"text":85,"@type":77},"It achieves larger gains in spatially constrained settings by learning adaptive modality selection, choosing whether to reason in vision or language at each step to improve both planning success and reasoning efficiency.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]