[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84607-en":3,"doc-seo-84607-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84607,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Reading Order Inference for Complex Document Layouts","Reading order inference addresses a key obstacle in digitizing complex historical manuscripts, where multiple reading streams interleave spatially and traditional layout assumptions break down. A training-free graph-based framework turns each OCR text line into a node and scores directed next-line transitions using lightweight language-model signals, then recovers the global order via a degree-constrained path cover. A max-regret rule reduces greedy edge-theft failures. Experiments on synthetic Glossa layouts, ALTO geometries, and OmniDocBench multi-column pages show large gains over recursive XY-cut and LayoutReader, including mirror invariance robustness.","arXiv :2607 .0 10 18v 1 [ cs .CL] 1 Jul 2026  \nReading Order Inference for Complex Document Layouts  \nIddo Hakim, Sharva Gogawale, Omer Ventura, Gal Grudka, Daria Vasyutinsky-Shapira, Berat Kurar-Barakat, and Nachum Dershowitz  \nSchool of Computer Science and AI, Tel Aviv University, Ramat Aviv, Israel {iddoh,sharvag,omerventura,[galgrudka](galgrudka}@mail.tau.ac.il)[}](galgrudka}@mail.tau.ac.il)[@mail.tau.ac.il](galgrudka}@mail.tau.ac.il)[ ](galgrudka}@mail.tau.ac.il){dariashap,berat,[nachumd](nachumd}@tauex.tau.ac.il)[}](nachumd}@tauex.tau.ac.il)[@tauex.tau.ac.il](nachumd}@tauex.tau.ac.il)  \nAbstract. Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple spatially interleaved reading streams, the canonical example being the Glossa Ordinaria layout, in which a central text is surrounded by commentaries that wrap around it in non-rectangular, non-convex regions. We present a training-free, graph-based framework: each OCR text line becomes a node in a directed candidate-transition graph, edges are scored by a weighted additive ensemble of two lightweight languagemodel signals (causal language model conditional likelihood and BERT next-sentence prediction, NSP; a third sentence-embedding signal was evaluated but did not improve reading order), and the global reading order is recovered as a degree-constrained directed path cover. To avoid the cascading “edge-theft” failures of greedy edge selection, we propose a max-regret inference rule that prioritizes commitments with high opportunity cost. We evaluate on synthetic Glossa Ordinaria grid layouts, on 23 ALTO page geometries (10 historical source pages plus mirrored and flipped variants), and on a 140-page multi-column English subset of OmniDocBench, comparing our method against the canonical recursive XY-cut (PaddleOCR PP-StructureV3) and two LayoutReader variants (layout-only and text+layout) on identical inputs. On wrap-around Glossa layouts our method recovers 95% of ground-truth successor edgeson average vs. XY-cut’s 50%; on the OmniDocBench multi-column subset it reaches 88% macro edge accuracy versus XY-cut’s 75% and LayoutReader’s 25% . The LayoutReader baselines transfer poorly due toa word-level vs. line-level granularity mismatch. We additionally verify mirror-invariance under horizontal and vertical page reflections: Our method changes by less than 1 percentage point, classical XY-cut by 2 points, and LayoutReader-T by up to 8 points.  \nKeywords: Reading order inference · Document layout analysis · NonManhattan layouts · Training-free · Max-regret inference · Historical manuscripts · XY-cut · LayoutReader · Mirror invariance  \nFig. 1: Two examples of non-Manhattan layouts. Left: a printed Hebrew Bible page, main text flanked by an Aramaic translation, two commentaries wrapping around them, Masorah parva as abbreviated notes in an internal margin, and Masorah magna spanning the top and bottom margins, all automatically (and imperfectly) line-segmented. Right: a page of a manuscript (Codex Bodmer 25) of the Greek Bible with two regions of catena commentary.  \n1 Introduction  \nReading order inference reconstructs the intended sequential flow of text elements detected by OCR. Accurate order is a prerequisite for searchable, continuous text streams and downstream document understanding; yet, it remains fragile whenever a page contains multiple spatially interleaved narratives.  \nFor simple single-column pages, ordering is recovered by a top-to-bottom scan. Visually complex documents often contain multiple partially independent reading streams: parallel columns, marginalia, interlinear glosses, side notes, and figure captions, where geometry alone is underdetermined: several plausible successors may be nearby in space but belong to different narrative tracks.  \nHistorical manuscripts amplify these difficulties. A canonical extreme is the Glossa Ordinaria page, where a central text is surroun","cbCainkDRLY6I85g","https://ap.wps.com/l/cbCainkDRLY6I85g","pdf",9773979,1,17,"English","en",105,"# Introduction\n## Goal and Key Idea\n## Graph-Based Training-Free Inference","[{\"question\":\"Why is reading order inference difficult for complex historical manuscript layouts?\",\"answer\":\"Because pages can contain multiple spatially interleaved reading streams, such as wrap-around commentaries in non-rectangular regions. Geometry alone becomes underdetermined, so several nearby successors may belong to different narrative tracks.\"},{\"question\":\"What is the core method proposed for recovering the reading order?\",\"answer\":\"The method builds a directed candidate-transition graph where each OCR text line is a node and edge scores reflect semantic continuity. It then recovers a globally consistent successor set using a degree-constrained directed path cover.\"},{\"question\":\"How does the approach avoid failures caused by greedy edge selection?\",\"answer\":\"It introduces a max-regret inference rule that prioritizes commitments with high opportunity cost. This is designed to prevent cascading “edge-theft” errors from local greedy choices.\"}]",1784197080,43,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"reading-order-inference-for-complex-document-layouts","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/reading-order-inference-for-complex-document-layouts/84607/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is reading order inference difficult for complex historical manuscript layouts?","Question",{"text":74,"@type":75},"Because pages can contain multiple spatially interleaved reading streams, such as wrap-around commentaries in non-rectangular regions. Geometry alone becomes underdetermined, so several nearby successors may belong to different narrative tracks.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What is the core method proposed for recovering the reading order?",{"text":79,"@type":75},"The method builds a directed candidate-transition graph where each OCR text line is a node and edge scores reflect semantic continuity. It then recovers a globally consistent successor set using a degree-constrained directed path cover.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the approach avoid failures caused by greedy edge selection?",{"text":83,"@type":75},"It introduces a max-regret inference rule that prioritizes commitments with high opportunity cost. This is designed to prevent cascading “edge-theft” errors from local greedy choices.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]