[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82710-en":3,"doc-seo-82710-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82710,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Alignment-Guided Largest Table Overlap Size Estimation","Fast estimation of the largest overlap size between tables enables effective blocking and query-by-table retrieval in large table repositories. The state-of-the-art estimator Armadillo improves efficiency by independently embedding each table and predicting overlap ratio via embedding similarity, but it struggles in heterogeneous settings. The proposed ALORE addresses three key issues: structural row–column alignment dependence, missing inter-table alignment signals, and value-encoding overfitting to corpus-specific distributions. ALORE instantiates row–column hypergraph modeling, alignment-guided training with inexpensive interactions, and domain-robust value mapping. Experiments on diverse datasets and a large real-world corpus show improved accuracy, up to 55% lower MAE, 69% in zero-shot transfer, and up to 89× speedup, with strong retrieval validation.","Alignment-Guided Largest Table Overlap Size Estimation  \nGe Lee∗ RMIT University Melbourne, Australia [ge.lee@student.rmit.edu.au](ge.lee@student.rmit.edu.au)  \nShixun Huang  \nUniversity of Wollongong Wollongong, Australia [shixunh@uow.edu.au](shixunh@uow.edu.au)  \nZhifeng Bao  \nThe University of Queensland Brisbane, Australia [zhifeng.bao@uq.edu.au](zhifeng.bao@uq.edu.au)  \nShazia Sadiq  \nThe University of Queensland Brisbane, Australia [s.sadiq@uq.edu.au](s.sadiq@uq.edu.au)  \nYanchang Zhao  \nData61, CSIRO Canberra, Australia [yanchang.zhao@csiro.au](yanchang.zhao@csiro.au)  \narXiv :2607 .03049v 1 [ cs .CL] 3 Jul 2026  \nAbstract  \nFast estimation of the size of the largest overlap between tables enables blocking and query-by-table retrieval in large table repositories. The first and the state-of-the-art estimator Armadillo improves efficiency by embedding each table independently and approximating overlap ratio via embedding similarity. However, accurate estimation in heterogeneous repositories remains limited by three challenges: (C1) overlap depends on row–column structure, i.e., each matched cell must preserve both its row and column membership under a joint alignment of the two tables, but existing encodings leave this structure to be inferred indirectly; (C2) independent encoding provides no explicit channel for inter-table alignment signals, biasing prediction toward global similarity; (C3) naïve value encodings overfit to corpus-specific distributions, causing cross-domain degradation. Hence, we propose ALORE, a scalable and domain-robust overlap ratio estimator built on three principles: (P1) explicitly represent row–column structure; (P2) expose inter-table alignment signals during training without expensive alignment search; (P3) reduce sensitivity to corpus-specific value distributions. ALORE instantiates these principles with a Two-View Row–Column Hypergraph encoder, alignment-guided objectives with inexpensive interaction signals, and a domain-robust value mapping. Experiments on multiple datasets spanning diverse domains and scales, including a large real-world corpus beyond prior benchmarks, show that ALORE outperforms the state of the art. ALORE reduces MAE by up to 55% overall and 69% in zero-shot transfer, while achieving up to 89× speedup. We further validate its effectiveness for query-by-table retrieval.  \n1 Introduction  \nWe study table overlap ratio estimation [43]: given two tables, estimate the size of their largest overlap. Intuitively, this is the number of cells in the largest common rectangular subtable that the two tables can share exactly in value. This subtable is obtained by reordering rows and columns through injective (one-to-one) row and column alignments. We refer to this size, normalized by the number of cells in the smaller table, as the overlap ratio.  \nConsider a workflow of repository search and curation over tables of property sales. A user provides a target table of property sale records and asks the system to find duplicate or near-duplicate tables in a large repository collected from public agencies, listing  \n∗ This work was done while the author was a visiting student at The University of Queensland and affiliated with Data61, CSIRO.  \nportals, and archived snapshots. Most candidates are unrelated, even though many may appear superficially similar. Relevant candidates may contain the same sale records with rows and columns reordered, cover only a subset of the target because they were extracted over different time windows, or have missing and inconsistent headers. These candidates are difficult to identify from metadata alone, and exact computation over all candidates is too expensive. The system therefore needs a fast filtering step that retains likely reordered copies and partial extracts, while discarding unrelated tables before expensive exact verification.  \nThe above illustrates why overlap ratio estimation is a useful primitive in large table repositories [43]. A fast est","cbCaipkPRy6iSEbj","https://ap.wps.com/l/cbCaipkPRy6iSEbj","pdf",3483818,3,1,17,"English","en",105,"# Introduction\n## Problem definition: largest common rectangular subtable\n## Use cases in large table repositories\n## Existing solutions and open challenges\n## Proposed ALORE approach (high-level)","[{\"question\":\"What does “largest overlap” mean in the table overlap ratio estimation task?\",\"answer\":\"It is the number of cells in the largest common rectangular subtable that two tables can share exactly by applying one-to-one injective alignments of rows and columns. The overlap ratio normalizes this size by the number of cells in the smaller table.\"},{\"question\":\"Why are existing estimators like Armadillo limited in heterogeneous table repositories?\",\"answer\":\"Their independent encoding and embedding-similarity prediction lacks explicit inter-table alignment signals, and their value encodings can overfit to corpus-specific distributions. Additionally, the row–column structure dependence is not modeled explicitly.\"},{\"question\":\"How does ALORE improve efficiency and accuracy compared with the state of the art?\",\"answer\":\"ALORE explicitly represents row–column structure, incorporates alignment-guided training using inexpensive interaction signals, and uses a domain-robust value mapping to reduce cross-domain degradation. Experiments show reduced MAE (up to 55% overall, 69% in zero-shot transfer), up to 89× speedup, and effectiveness for query-by-table retrieval.\"}]",1784182435,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"alignment-guided-largest-table-overlap-size-estimation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/alignment-guided-largest-table-overlap-size-estimation/82710/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does “largest overlap” mean in the table overlap ratio estimation task?","Question",{"text":75,"@type":76},"It is the number of cells in the largest common rectangular subtable that two tables can share exactly by applying one-to-one injective alignments of rows and columns. The overlap ratio normalizes this size by the number of cells in the smaller table.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why are existing estimators like Armadillo limited in heterogeneous table repositories?",{"text":80,"@type":76},"Their independent encoding and embedding-similarity prediction lacks explicit inter-table alignment signals, and their value encodings can overfit to corpus-specific distributions. Additionally, the row–column structure dependence is not modeled explicitly.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ALORE improve efficiency and accuracy compared with the state of the art?",{"text":84,"@type":76},"ALORE explicitly represents row–column structure, incorporates alignment-guided training using inexpensive interaction signals, and uses a domain-robust value mapping to reduce cross-domain degradation. Experiments show reduced MAE (up to 55% overall, 69% in zero-shot transfer), up to 89× speedup, and effectiveness for query-by-table retrieval.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]