[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120209-en":3,"doc-seo-120209-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120209,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Debiased Gaussian Process-based Machine Learning with Partially Observed Information - Two-Stage Model","Widely applicable machine learning and artificial intelligence technologies face increasing demand for reliable models under real-world data scarcity. Although data augmentation helps, bias remains unavoidable and can degrade prediction accuracy. A Two-Stage Debiased Gaussian Process (TSDGP) framework is proposed to deliver robust, accurate predictions with partially observed information. It uses a latent-variable reconstruction in stage one, then refines the model and uncertainty via Bayesian updating in stage two on augmented data, supported by theoretical and experimental evidence.","Proceedings of the 58th Hawaii International Conference on System Sciences | 2025  \nDebiased Gaussian Process-based Machine Learning with Partially  \nObserved Information  \nYanwen Xu University of Texas at Dallas  \n [yanwen.xu@utdallas.edu](yanwen.xu@utdallas.edu)  \nShengxiang Wu JP Morgan Chase & CO.  [shengxiang.wu@chase.com](shengxiang.wu@chase.com)  \nXuehui Chao JP Morgan Chase & CO.  [xuehui.chao@chase.com](xuehui.chao@chase.com)  \nAbstract  \nWidely applicable machine learning and artificial intelligence technologies have resulted in an increasing demand for reliable models. Due to the ubiquitous data scarcity in the real world, model training can often be challenging and face limitations. Although various data augmentation techniques can efficiently alleviate this dilemma, bias is unavoidable, causing trade-offs in prediction accuracy. As this general dilemma is addressed, we discuss a Two-Stage Debiased Gaussian Process (TSDGP)-based machine learning model capable of providing robust and accurate predictions across various fields, even with partially observed information. Given the partially observed information in input data, the latent variable model was leveraged to enhance heterogeneous data utilization by reconstructing the unavailable information in stage one. Subsequently, the model and uncertainties from the first stage were refined within the Bayesian framework using the augmented dataset in stage two. By demonstrating the consistency and first and second moments of the proposed two-stage model, we are confident in the accuracy and robustness of the results. Supported by solid theoretical proof, we further evaluate the results of TSDGP through numerical and empirical experiments, showing the premium performances of the proposed approach. In conclusion, TSDGP can solve the dilemma caused by data scarcity in the real world—enabling a reliable high-fidelity predictive model to be trained on partially observed datasets without a significant trade-off in accuracy.  \nKeywords: Gaussian Process, Latent Variable, Data Augmentation, Debiased Machine Learning, Partially Observed Information, Uncertainty Propagation  \n1. Introduction  \nNowadays, Machine Learning (ML) and Artificial Intelligence (AI) applications are blooming as revolutionary algorithms developed and more accessible computing power. Estimators are trained to achieve high accuracy in a wide range of fields, including image recognition, natural language processing, autonomous driving, and predictive analytics. Reliable estimators, especially for Large Language Model (LLM) and diffusion models (DM), are often trained with massive amounts of labeled data. For instance, GPT-3 was trained based on 175 billion parameters in the pre-training phase, and trained with specific labeled data and reinforcement learning from human feedback (RLHF) .  \nHowever, the robust data and massive engineering that gave birth to GPT-3 cannot easily be replicated. In the real world, data usually appears to be scarce and noisy. Without sufficient labeled data, training reliable estimators can be difficult. Fortunately, remarkable achievements in data augmentation techniques efficiently alleviated the high demand of data. By randomly replacing words in a sentence with other words to increase training samples while preserving the original sentence structure, X. Wanget al. (2018) applied ”SwitchOut” technique to improve the variability and robustness during the training process. Wei and Zou (2019) improved the performance of text classification models with limited amount of data by through applying synonym replacement, random insertion, random swap, and random deletion. Such technique efficiently increased the training data size. Similarly, Xie et al. (2020) applied back-translation to improve the performance of language models in semi-supervised learning with little available data.  \nData augmentation is the process of representing the  \nURI: [https://hdl.handle.net/10125/108942](https:","cbCaisN0dA4eCzXx","https://ap.wps.com/l/cbCaisN0dA4eCzXx","pdf",1290808,1,10,"English","en",105,"# Introduction\n## Data scarcity and bias in real-world machine learning\n## Data augmentation and imputation background\n## Overview of the proposed TSDGP framework","[{\"question\":\"Why is training reliable machine learning models difficult with partially observed information?\",\"answer\":\"Real-world data is often scarce and noisy, and insufficient labeled data makes estimator training unreliable. Partially observed inputs further limit accurate learning and can lead to biased predictions.\"},{\"question\":\"How does TSDGP use partially observed information during modeling?\",\"answer\":\"Stage one leverages a latent variable model to reconstruct unavailable information and enhance heterogeneous data utilization. Stage two refines both the model and uncertainties using Bayesian updating with the augmented dataset.\"},{\"question\":\"What evidence supports the accuracy and robustness of the proposed approach?\",\"answer\":\"The paper demonstrates consistency and the first and second moments of the two-stage model, supported by theoretical proof. It also reports numerical and empirical experiments showing superior performance compared with alternatives.\"}]","Debiased Gaussian Process-based Machine Learning with Partially Observed Information - Two-Stage Model | PDF",1785728740,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"debiased-gaussian-process-based-machine-learning-with-partially-observed-information-two-stage-model","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/debiased-gaussian-process-based-machine-learning-with-partially-observed-information-two-stage-model/120209/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is training reliable machine learning models difficult with partially observed information?","Question",{"text":75,"@type":76},"Real-world data is often scarce and noisy, and insufficient labeled data makes estimator training unreliable. Partially observed inputs further limit accurate learning and can lead to biased predictions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TSDGP use partially observed information during modeling?",{"text":80,"@type":76},"Stage one leverages a latent variable model to reconstruct unavailable information and enhance heterogeneous data utilization. Stage two refines both the model and uncertainties using Bayesian updating with the augmented dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence supports the accuracy and robustness of the proposed approach?",{"text":84,"@type":76},"The paper demonstrates consistency and the first and second moments of the two-stage model, supported by theoretical proof. It also reports numerical and empirical experiments showing superior performance compared with alternatives.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]