[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128746-en":3,"doc-seo-128746-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128746,1099523885074,"Ivy","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Improving Deep Image Feature Matching","Image matching establishes pixel correspondences between two images by extracting local features and matching them via similarity. Matching is studied in two sub-domains: geometric matching that aligns pixels from the same 3D point, and semantic matching that links pixels with consistent meaning. The thesis improves geometric accuracy with a dual-resolution 4D convolution framework and simplifies semantic matching via a metric-learning-inspired training setup. It further explores learning 3D attributes from 2D image features for pose, shape, and garment disentanglement.","Improving Deep Image Feature  \nMatching  \nXinghui Li  \nLincoln College University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy Michaelmas 2025  \n“We can’t solve problems by using the same kind of thinking we used when we  \ncreated them.” – Albert Einstein  \nAbstract  \nImage matching establishes pixel correspondences between two images. It is usually achieved in two steps, where the high-dimensional local image features are first extracted across images and correspondences are established based on features’similarities. Image matching can be categorized into two sub-domains: geometric and semantic matching, where the former matches pixels describing the same 3D point while the latter finds pixels having the same semantic meaning. This thesis focuses on advancing both types of matching from various perspectives and explores another application of image features; namely learning 3D attributes from 2D images.  \nFor geometric matching, we propose improvements in accuracy through refined network architectures. NCNet by Rocco et al. [1], demonstrates that 4D convolution can be used to filter incorrect matches and improve matching accuracy. However, the quadratic complexity of this module makes it impractical for matching high-resolution images, a potentially crucial factor in achieving accurate geometric matching. To address this limitation, we propose a dual-resolution network architecture. Our method applies 4D convolution to filter matches on a coarse scale and uses this result to guide the matching on a finer scale. This approach significantly enhances accuracy by increasing the matching resolution while maintaining low computational overhead.  \nFor semantic matching, we first propose a simple and effective training framework that significantly reduces model complexity. Previous works typically follow connecting a pre-trained feature backbone with a complex matching module to handle appearance differences between semantically similar points. Inspired by metric learning, we show that the feature backbone itself can achieve comparable performance without any additional module, hence greatly simplifying the overall design. Building on this framework, we next explore adapting large vision foundation models for semantic matching. Earlier works found that Stable Diffusion (SD), known for its text-to-image generation capabilities, is also able to produce highly discriminative features for image matching. We demonstrate that the matching capability of SD can be enhanced through prompt tuning, and propose a conditional prompting module that infuses prior knowledge of the image pair as the SD’s prompt to further improve matching accuracy.  \nBeyond feature matching, we investigate inferring 3D information from 2D images, a different use case of image features. Previous studies have shown that a  \n3D model of a dressed person can be derived from a 2D image feature map. We found that a person’s pose, shape, and garment attributes can be disentangled from this feature map. These attributes can be swapped or replaced, allowing the creation of a new 3D model from the modified feature map. Unlike typical disentanglement tasks like face swapping, which operate in the same data domain, our work transitions from 2D to 3D and is able to generate 3D models from 2D image attributes.  \nTo summarise, this thesis enhances both geometric and semantic matching, and also advances the inference of 3D information from 2D image features.  \nKeywords – Image Feature Matching, Semantic Matching, Geometric Matching, Diffusion Model, Attribute Disentanglement.  \nDeclaration  \nThis thesis is submitted to the Department of Engineering Science, University of Oxford, in fulfilment of the requirements for the degree of Doctor of Philosophy. This thesis is entirely my own work, and except where otherwise stated, describes my own research.  \nXinghui Li, December 2024 .  \nAcknowledgment  \nThis is it—I have reached the destination I set out for five ","cbCaifqEkcZpnkiQ","https://ap.wps.com/l/cbCaifqEkcZpnkiQ","pdf",18289665,1,124,"English","en",105,"# 1 Introduction\n## 1.1 Key","[{\"question\":\"What problem does the thesis address in image matching?\",\"answer\":\"It addresses how to establish accurate correspondences between two images by improving both geometric and semantic matching based on learned image features.\"},{\"question\":\"How does the proposed approach improve geometric feature matching?\",\"answer\":\"It uses a dual-resolution network that applies 4D convolution at a coarse scale to filter incorrect matches, then guides fine-scale matching to increase accuracy with manageable computation.\"},{\"question\":\"What additional application beyond matching does the thesis investigate?\",\"answer\":\"It investigates inferring 3D information from 2D images, showing that pose, shape, and garment attributes can be disentangled from 2D feature maps and then used to generate new 3D models.\"}]","Improving Deep Image Feature Matching | PDF",1786003068,312,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"improving-deep-image-feature-matching","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-deep-image-feature-matching/128746/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address in image matching?","Question",{"text":76,"@type":77},"It addresses how to establish accurate correspondences between two images by improving both geometric and semantic matching based on learned image features.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed approach improve geometric feature matching?",{"text":81,"@type":77},"It uses a dual-resolution network that applies 4D convolution at a coarse scale to filter incorrect matches, then guides fine-scale matching to increase accuracy with manageable computation.",{"name":83,"@type":74,"acceptedAnswer":84},"What additional application beyond matching does the thesis investigate?",{"text":85,"@type":77},"It investigates inferring 3D information from 2D images, showing that pose, shape, and garment attributes can be disentangled from 2D feature maps and then used to generate new 3D models.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]