[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-133250-en":3,"doc-seo-133250-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},133250,962084925636,"Sophia Brooks","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","InstructAny2Pix - Image Editing with Multi-Modal Prompts - Multi-Object Instruction System","InstructAny2Pix advances instruction-following image editing by using a multi-modal LLM to execute complex edit instructions. It supports multi-object edits, interleaving text instructions with multiple reference images, and even audio or music inputs for creative editing scenarios. To measure performance, the work introduces two new multimodal benchmark datasets, MM-Inst and Dreambooth++, and evaluates the model against multimodal and conventional image editing baselines.","InstructAny2Pix: Image Editing with Multi-Modal prompts  \nShufan Li 1 , Harkanwar Singh1 , Aditya Grover1  \n1University of California, Los Angeles  \nCorrespondence: [jacklishufan@cs.ucla.edu](jacklishufan@cs.ucla.edu)  \nAbstract  \nImage editing has made incredible progress in recent years. Early works only supported caption-guided editing, but recently, free-form text instructions and reference images have been incorporated to allow for more flexibility.  \nHowever, existing methods still struggle with complex editing instructions involving multiple objects or reference images. We present InstructAny2Pix, a novel image editing model that leverages a multi-modal LLM to execute intricate edit instructions. Compared with previous works, InstructAny2Pix extends the flexibility of edit instructions in three key ways:  \nFirst, it can perform complex instructions involving multiple object edits; second, it supports the interleaving of text instructions with multiple reference images; and third, it supports audio and music inputs as part of the edit prompts, unlocking creative applications such as album cover generation and music-inspired merchandise design. To evaluate the effectiveness of InstructAny2Pix, we propose two new benchmark datasets, MM-Inst and Dreambooth++, consisting of human-written, multimodal prompts. InstructAny2Pix outperforms baselines on these two proposed multi-modal benchmarks, as well as on conventional image editing benchmarks such as InstructPix2Pix.  \n1 Introduction  \nThe ability to edit an existing image using freeform text instructions vastly expands the usability of image editing models. Compared with early works, such as Prompt2Prompt(Hertz et al., 2022), which require caption pairs, instructionbased image editing methods, such as InstructPix2Pix(Brooks et al., 2023), offer users unparalleled flexibility to describe edit instructions in natural language, such as \"add a dog.\" More recent models, such as Kosmos-G(Pan et al., 2023), additionally accept reference images, allowing users to add a specific dog from the reference image to the  \nscene. Despite these progresses, existing methods still have limited instruction-following capabilities. For text-guided edits, they are limited to simple instructions on which they were trained and cannot generalize to complex instructions involving multiple objects, such as \"add a wolf howling under the moon.\" For image-guided edits, they often struggle to complete complex instructions, such as \"replace the cat with [reference image],\" while faithfully respecting both the image to edit and the reference image.  \nTo address these limitations, we propose InstructAny2Pix, the first instruction-following image editing system capable of following a wide range of complex, multi-modal, multi-object instructions. Specifically, InstructAny2Pix not only supports text instructions involving multiple objects, such as\"add a wolf howling under the moon\" or \"add a cat and remove the dog,\" but it can also optionally accept multiple reference images of the objects (e.g., the wolf and the moon) . Furthermore, it works with arbitrary free-form, multi-modal instructions interleaving text, image, and audio, such as \"change [image A] to the style of [image B]\" or \"fit [image] to [music],\" while previous multi-modal models only support limited modalities (i.e., image) and very basic instructions (e.g., add, remove) .  \nInstructAny2Pix greatly enhances the flexibility and usability of image editing models. When creating a scene with multiple objects, instead of writing lengthy descriptions for each object, uploading reference images can be far more efficient. Music or audio inputs, though less obvious, also unlock creative possibilities, such as designing T-shirts based on music or dynamically adapting a background image during live performances. While these tasks could be done through text instructions, they would require designers to first develop specific ideas like\"add a circle to the T-sh","cbCaiuL6XYp3cwdv","https://ap.wps.com/l/cbCaiuL6XYp3cwdv","pdf",20171144,1,26,"English","en",105,"# Abstract\n# Introduction\n## Text-guided and image-guided editing limitations\n## Proposed approach: InstructAny2Pix capabilities\n## Training data construction and model architecture","[{\"question\":\"What does InstructAny2Pix enable for image editing?\",\"answer\":\"It enables instruction-following image editing driven by multi-modal prompts, including complex multi-object instructions and optional multiple reference images. It can also incorporate audio or music inputs as part of the edit prompt.\"},{\"question\":\"How is InstructAny2Pix different from earlier text- or image-guided editing methods?\",\"answer\":\"Earlier text-guided methods often generalize only to simple training patterns, while image-guided methods struggle to satisfy both the edited image and the reference constraints for complex instructions. InstructAny2Pix is designed to follow a broader range of complex, multi-modal, multi-object instructions.\"},{\"question\":\"What benchmarks are proposed to evaluate the model?\",\"answer\":\"The paper proposes two benchmark datasets, MM-Inst and Dreambooth++, built from human-written multimodal prompts. The model is reported to outperform baselines on these datasets and on conventional image editing benchmarks such as InstructPix2Pix.\"}]","InstructAny2Pix - Image Editing with Multi-Modal Prompts - Multi-Object Instruction System | PDF",1787216349,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"instructany2pix-image-editing-with-multi-modal-prompts-multi-object-instruction-system","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/instructany2pix-image-editing-with-multi-modal-prompts-multi-object-instruction-system/133250/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-25","2026-08-20",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does InstructAny2Pix enable for image editing?","Question",{"text":76,"@type":77},"It enables instruction-following image editing driven by multi-modal prompts, including complex multi-object instructions and optional multiple reference images. It can also incorporate audio or music inputs as part of the edit prompt.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is InstructAny2Pix different from earlier text- or image-guided editing methods?",{"text":81,"@type":77},"Earlier text-guided methods often generalize only to simple training patterns, while image-guided methods struggle to satisfy both the edited image and the reference constraints for complex instructions. InstructAny2Pix is designed to follow a broader range of complex, multi-modal, multi-object instructions.",{"name":83,"@type":74,"acceptedAnswer":84},"What benchmarks are proposed to evaluate the model?",{"text":85,"@type":77},"The paper proposes two benchmark datasets, MM-Inst and Dreambooth++, built from human-written multimodal prompts. The model is reported to outperform baselines on these datasets and on conventional image editing benchmarks such as InstructPix2Pix.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]