[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85327-en":3,"doc-seo-85327-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85327,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Direct Image to Modern Vietnamese Translation of Han Nom Manuscripts via Multimodal RLHF Preference Alignment","Translating Hán–Nôm manuscripts into modern Vietnamese is challenging due to degraded historical pages, rare logographic characters, and limited parallel supervision. A multimodal RLHF preference-alignment framework conditions Vietnamese generation on manuscript images and aligned Hán–Nôm source text. The model fuses visual CLIP features, Hán–Nôm embeddings, Vietnamese representations, and T5 states into a shared latent space, then compares PPO, DPO, and KTO. DPO achieves the strongest lexical, semantic, and token-level quality, while preference alignment improves over SFT for low-resource historical translation.","DIRECT IMAGE-TO-MODERN VIETNAMESE TRANSLATION OF HAN-NOM MANUSCRIPTS VIA MULTIMODAL RLHF PREFERENCE ALIGNMENT  \nThi Kim Trang Vo 1,2 , Nghia Hieu Nguyen 1,2, Ha Minh Tan 1,2  \n1 University of Information Technology, Ho Chi Minh City, Vietnam  \n2 Vietnam National University, Ho Chi Minh City, Vietnam [trangvtk.18@grad.uit.edu.vn](trangvtk.18@grad.uit.edu.vn), [nghiangh@uit.edu.vn](nghiangh@uit.edu.vn), [tanhm@uit.edu.vn](tanhm@uit.edu.vn).  \narXiv :2607 . 1 1434v 1 [ cs .CL] 13 Jul 2026  \nABSTRACT  \nTranslating Hán–Nôm manuscripts into modern Vietnamese is challenging because historical pages are degraded, the script contains rare logographic characters, and parallel supervision is limited. We propose a multimodal RLHF preferencealignment framework that conditions Vietnamese generation on manuscript images and aligned Hán–Nôm source text. The model combines four streams: CLIP ViT-L/14@336 for visual features, bert-base-chinese for Hán–Nôm representations, vinai/phobert-base for Vietnamese representations, and T5-small encoder states. Modality-specific projections and a fusion block compress the resulting 2,048-dimensional concatenation into a shared 512-dimensional representation. Starting from the same supervised fine-tuned policy, we compare PPO, DPO, and KTO under matched work-level macro-averaged evaluation. DPO achieves the best BLEU-4, ROUGE-L, BERTScore, semantic similarity, CER, WER, and token accuracy, whereas PPO obtains the highest precision, recall, and F1 . KTO remains competitive through its desirable–undesirable utility objective. All preference-aligned policies improve the BLEU-4 and semantic-similarity scores available for the SFT baseline. These results indicate that multimodal preference optimization complements supervised learning for lexical and semantic quality in low-resource historical translation.  \nIndex Terms— Historical Han-Nom document translation, Multimodal Fusion, Reinforcement Learning with Human Feedback (RLHF), Reward/Critic Modeling, Preference-based alignment, PPO/DPO/KTO algorithm.  \n1 Introduction  \nHistorical Vietnamese manuscripts written in Hán–Nôm preserve important cultural, literary, and administrative knowledge, but remain difficult to access. The script contains rare logographic characters, and surviving pages often suffer from degradation, ink diffusion, irregular handwriting, and layout variation. Parallel data linking manuscript images, Hán–Nôm transcriptions, and fluent modern Vietnamese translations are scarce. The task is therefore not only OCR or text-to-text translation, but a coupled image-to-language problem: the system  \nmust read degraded manuscript evidence, recover source-script content, and generate faithful modern Vietnamese.  \nPrior work has mostly addressed this pipeline in separate stages. OCR resources such as NomNaOCR [1], IHRNomDB [2], and the Nom–Vietnamese Parallel Corpus [3] support recognition, transliteration, and translation research, while OCR systems such as PaddleOCRv5 [4] improve robustness on degraded pages. However, OCR maps images to Hán– Nôm text and does not directly generate fluent Vietnamese. In parallel, SMT, NMT, and Transformer methods have been used for Nôm-to-Quốc-ngữ, Chinese–Vietnamese, and Hán–Nôm– Vietnamese translation [5, 6, 7, 8], but these methods usually assume clean source text. LLM-based post-OCR correction [9] reduces recognition noise, yet still operates after OCR and does not jointly align visual evidence, source-script semantics, and target-language fluency.  \nPure supervised Seq2Seq training is also limited for classical and literary Hán–Nôm. A passage may have multiple valid Vietnamese renderings: one may be literal but unnatural, while another better preserves meaning, rhythm, and literary nuance. Maximum-likelihood training imitates a single reference and does not model preferences among alternatives, while BLEU-like metrics may miss fluency, cultural adequacy, and style [10, 11] . These issues motivate a preference-bas","cbCaipNKALEQvKKF","https://ap.wps.com/l/cbCaipNKALEQvKKF","pdf",2115617,9,1,6,"English","en",105,"# Abstract\n# Introduction\n## Challenges in Hán–Nôm Translation\n## Limits of Prior Pipelines and Supervised Learning\n# Methodology\n## Data and Preference Construction\n## Preference Optimization Algorithms","[{\"question\":\"Why is translating Hán–Nôm manuscripts into modern Vietnamese difficult?\",\"answer\":\"Historical pages are often degraded, the script contains rare logographic characters, and aligned parallel data linking images, Hán–Nôm transcriptions, and fluent Vietnamese is scarce.\"},{\"question\":\"How does the proposed system generate Vietnamese directly from manuscript images?\",\"answer\":\"It learns a multimodal policy that uses manuscript image evidence together with training-time Hán–Nôm supervision and aligned source text representations, avoiding OCR-decoder-dependent inference.\"},{\"question\":\"What differences exist among PPO, DPO, and KTO in this work and which performs best?\",\"answer\":\"PPO uses reward/critic learning, DPO optimizes chosen–rejected pairs directly, and KTO uses a utility-based objective robust to noisy rejected samples. DPO yields the best overall scores (BLEU-4, ROUGE-L, BERTScore, similarity, CER/WER, and token accuracy).\"}]",1784202517,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"direct-image-to-modern-vietnamese-translation-of-han-nom-manuscripts-via-multimodal-rlhf-preference-alignment","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/direct-image-to-modern-vietnamese-translation-of-han-nom-manuscripts-via-multimodal-rlhf-preference-alignment/85327/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is translating Hán–Nôm manuscripts into modern Vietnamese difficult?","Question",{"text":76,"@type":77},"Historical pages are often degraded, the script contains rare logographic characters, and aligned parallel data linking images, Hán–Nôm transcriptions, and fluent Vietnamese is scarce.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed system generate Vietnamese directly from manuscript images?",{"text":81,"@type":77},"It learns a multimodal policy that uses manuscript image evidence together with training-time Hán–Nôm supervision and aligned source text representations, avoiding OCR-decoder-dependent inference.",{"name":83,"@type":74,"acceptedAnswer":84},"What differences exist among PPO, DPO, and KTO in this work and which performs best?",{"text":85,"@type":77},"PPO uses reward/critic learning, DPO optimizes chosen–rejected pairs directly, and KTO uses a utility-based objective robust to noisy rejected samples. DPO yields the best overall scores (BLEU-4, ROUGE-L, BERTScore, similarity, CER/WER, and token accuracy).","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]