[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86187-en":3,"doc-seo-86187-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86187,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Enhancing LLMs through Human Feedback: A Journey Towards Self-Improvement","In the rapidly evolving landscape of information retrieval systems, the study presents a methodology to refine a primary Retrieval Augmented Generation (RAG) system by integrating an auxiliary feedback RAG system. Human-in-the-loop feedback is continuously collected, classified, and inserted into the inference workflow to enable iterative learning. Experiments evaluate performance on three benchmark datasets for general and custom domain knowledge using an LLM-as-a-Judge evaluation, showing improved accuracy, relevance, and response quality.","Enhancing LLMs through human feedback: a journey towards self-improvement  \nTatiana Pelc1 , Gila Kamhi1 , Asaf Avrahamy1 and Adi Fledel-Alon1, *  \n1Intel Corporation  \nAbstract  \nIn the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. This study introduces a novel methodology for refining the performance of a primary Retrieval Augmented Generation (RAG) system by strategically integrating an auxiliary feedback RAG system. By systematically harnessing human-generated feedback, the approach aims to enhance the accuracy, relevance, and overall quality of responses, driving the system towards self-improvement. Central to this methodology is a human-in-the-loop implementation, where user feedback is continuously collected, classified, and integrated into the inference workflow, enabling the system to learn and evolve iteratively. To validate the effectiveness of this approach, the study employs rigorous testing against three diverse benchmark datasets focused on general and custom domain knowledge, utilizing a LLM-as-a-Judge evaluation strategy. This comprehensive framework not only underscores the transformative potential of feedback-driven enhancements in RAG systems but also sets a precedent for future research in adaptive information retrieval technologies, marking a significant step in the journey towards autonomous refinement and optimization through user engagement.  \nKeywords  \nRLHF, Retrieval Augmented Generation, Large language models, LLM-as-a-Judge  \n1. Introduction  \nLarge Language Models (LLMs) have driven recent advancesin AI, enabled by their unprecedented scale and improved through Reinforcement Learning from Human Feedback (RLHF) . RLHF refines model behavior using human input, significantly enhancing adaptability, responsiveness, and overall performance.  \nHowever, traditional AI systems have relied on binary feedback (e.g., like/dislike), which limits learning depth. This paper proposes incorporating rich textual human feedback—including detailed corrections—into the generative AI pipeline. This enables more nuanced learning and improves response accuracy and contextual alignment.  \n1.1. Related work  \nRetrieval Augmented Generation (RAG) enhances LLMs by retrieving relevant documents, but often fails to capture user preferences, resolve contextual mismatches, or adapt to evolving information needs. Common issues include retrieval bias, robustness challenges, and fairness and privacy concerns [1, 2] . Early approaches with static retrieval and binary feedback offered limited gains. Recent work emphasizes the role of structured and dynamic feedback to improve alignment and reliability. Hybrid methods combining human judgment with automated evaluation have been explored to reduce hallucinations. Human-in-the-loop systems and cross-domain adaptation are emerging as effective strategies for building self-improving, context-aware RAG systems.  \n1.1.1. Feedback in RAG and LLM systems  \nStructured feedback pipelines such as CDF-RAG (Causal Dynamic Feedback for Adaptive Retrieval-Augmented Generation) refine retrieval and generation through iterative loops guided by causal graphs, enabling multi-hop reasoning and improving accuracy [3] . In educational applications,  \n* Corresponding author.  \n$ [adi.fledel.alon@intel.com](adi.fledel.alon@intel.com) (A. Fledel-Alon)  \n© 2025 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4 .0 International (CC BY 4 .0) .  \ntimestamp-anchored feedback has reduced hallucinations by 37% and improved user satisfaction and learning outcomes [4] .  \n1.1.2. Cross-domain, comparative, multi-modal feedback integration  \nLLM-Cure addresses large-scale user review analysis by combining automated feature extraction with competitive analysis [5] . Analyzing over one million reviews from 70 apps, it introduces a three-stage remediation process. The method achieves an 85%","cbCaimL4gnWQFun0","https://ap.wps.com/l/cbCaimL4gnWQFun0","pdf",1664503,3,1,7,"English","en",105,"# Abstract\n# Introduction\n## Related work\n### Feedback in RAG and LLM systems\n### Cross-domain, comparative, multi-modal feedback integration\n### Feedback driven alignment\n# Our contribution","[{\"question\":\"How does the proposed method improve a primary RAG system using feedback?\",\"answer\":\"It integrates an auxiliary feedback RAG system and uses a human-in-the-loop process to collect, classify, and embed user feedback into the inference workflow for iterative refinement.\"},{\"question\":\"What limitations of traditional feedback (e.g., like/dislike) does the paper address?\",\"answer\":\"Binary feedback is too shallow for effective learning, so the approach leverages richer textual feedback such as detailed corrections to improve alignment with user intent and contextual accuracy.\"},{\"question\":\"How is the effectiveness of the approach evaluated?\",\"answer\":\"The study tests on three diverse benchmark datasets and applies an LLM-as-a-Judge evaluation strategy to measure improvements in answer quality, accuracy, and relevance.\"}]",1784209237,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"enhancing-llms-through-human-feedback-a-journey-towards-self-improvement","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/enhancing-llms-through-human-feedback-a-journey-towards-self-improvement/86187/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the proposed method improve a primary RAG system using feedback?","Question",{"text":75,"@type":76},"It integrates an auxiliary feedback RAG system and uses a human-in-the-loop process to collect, classify, and embed user feedback into the inference workflow for iterative refinement.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitations of traditional feedback (e.g., like/dislike) does the paper address?",{"text":80,"@type":76},"Binary feedback is too shallow for effective learning, so the approach leverages richer textual feedback such as detailed corrections to improve alignment with user intent and contextual accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the effectiveness of the approach evaluated?",{"text":84,"@type":76},"The study tests on three diverse benchmark datasets and applies an LLM-as-a-Judge evaluation strategy to measure improvements in answer quality, accuracy, and relevance.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]