[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118782-en":3,"doc-seo-118782-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118782,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","RRAML: Reinforced Retrieval Augmented Machine Learning","Large language models (LLMs) have transformed machine learning by enabling strong capabilities in understanding, generating, and manipulating human language. Conventional API-based prompting restricts context handling and makes external knowledge access difficult, while also increasing hallucination risk. Existing retrieval-augmented and memory-based methods rely on embedding models and often require adapting embeddings or access to LLM parameters. RRAML integrates LLM reasoning with supporting evidence retrieved from a user database via a purpose-built retriever, using reinforcement learning to avoid LLM gradients, reduce retraining needs, align retrieval with reasoning, and mitigate hallucinations and irrelevant or harmful documents.","arXiv :2307 . 12798v 3 [ cs .CL] 27 Jul 2023  \nRRAML: Reinforced Retrieval Augmented Machine Learning  \nAndrea Bacciu 1 , Florin Cuconasu 1 , Federico Siciliano 1 , Fabrizio Silvestri 1 , Nicola Tonellotto2 , and Giovanni Trappolini 1  \n[f](fsurnameg@diag.uniroma1.it1)[surname](fsurnameg@diag.uniroma1.it1)[g](fsurnameg@diag.uniroma1.it1)[@diag.uniroma1.it](fsurnameg@diag.uniroma1.it1)[1](fsurnameg@diag.uniroma1.it1) , [nicola.tonellotto@unipi.it](nicola.tonellotto@unipi.it2)[2](nicola.tonellotto@unipi.it2)  \n1 Sapienza University of Rome  \n2 University of Pisa  \nAbstract. The emergence of large language models (LLMs) has revolutionized machine learning and related 􀀌elds, showcasing remarkable abilities in comprehending, generating, and manipulating human language.  \nHowever, their conventional usage through API-based text prompt submissions imposes certain limitations in terms of context constraints and external source availability. LLMs su􀀋er from the problem of hallucinating text, and in the last year, several approaches have been devised to overcome this issue: adding an external Knowledge Base or an external memory consisting of embeddings stored and retrieved by vector databases. In all the current approaches, though, the main issues are:  \n(i) they need to access an embedding model and then adapt it to the task they have to solve; (ii) in case they have to optimize the embedding model, they need to have access to the parameters of the LLM, which in many cases are \\black boxes\". To address these challenges, we propose a novel framework called Reinforced Retrieval Augmented Machine Learning (RRAML) . RRAML integrates the reasoning capabilities of LLMs with supporting information retrieved by a purpose-built retriever from a vast user-provided database. By leveraging recent advancements in reinforcement learning, our method e􀀋ectively addresses several critical challenges. Firstly, it circumvents the need for accessing LLM gradients. Secondly, our method alleviates the burden of retraining LLMs for speci􀀌c tasks, as it is often impractical or impossible due to restricted access to the model and the computational intensity involved.  \nAdditionally, we seamlessly link the retriever's task with the reasoner, mitigating hallucinations and reducing irrelevant and potentially damaging retrieved documents. We believe that the research agenda outlined in this paper has the potential to profoundly impact the 􀀌eld of AI, democratizing access to and utilization of LLMs for a wide range of entities.  \nKeywords: Deep Learning · Information Retrieval · Large Language Models.  \n1 Introduction  \nThe advent of Large Language Models (LLMs) has brought about a paradigm shift in machine learning and its related disciplines. LLMs [2,20,13,24,1] have  \n2 Bacciu, Cuconasu, Siciliano, Silvestri, Tonellotto, Trappolini, 2023  \nexhibited unprecedented capabilities in understanding, generating, and manipulating the human language. Famously, ChatGPT [13] has entered the public space by reaching one million users in a matter of days. The way these models are usually used is through API that only allows submitting a textual prompt and getting back from the server the generated text. However, this causes an immediate limitation: all information must be passed through this context, and we know transformer-based models do not scale nicely. Even if they did, API costs are charged on the basis of their usage. Therefore, using long contexts would be expensive. Even if one had the resources to run their own LLM, the costs of training and of the hardware infrastructure, and the environmental impact should be considered. There is an impendent need, though, to accommodate the enormous power of those models to speci􀀌c user needs by making sure that they could use the reasoning capabilities of LLMs, through in-context learning [2] on their data.  \nA solution is to adopt a retrieval-augmented approach [8,26] . In this setting, a retriever is used to 􀀌lter out releva","cbCaiaYXWt2hYLfB","https://ap.wps.com/l/cbCaiaYXWt2hYLfB","pdf",274220,1,9,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limits of API prompting\n## Retrieval-augmented approaches and misalignment\n## Proposed solution: Reinforced Retrieval Augmented Machine Learning (RRAML)","[{\"question\":\"What problem does RRAML address in retrieval-augmented LLM usage?\",\"answer\":\"RRAML targets the misalignment between retriever and reasoner, which can lead to irrelevant or even dangerous retrieved content and consequently hallucinations.\"},{\"question\":\"How does RRAML reduce reliance on retraining or gradient access to LLMs?\",\"answer\":\"RRAML uses reinforcement learning to improve the system without needing access to LLM gradients, and it alleviates the need to retrain LLMs for specific tasks.\"},{\"question\":\"What is the core design of RRAML?\",\"answer\":\"RRAML couples an efficient retriever that searches a large user-provided database with a reasoner (e.g., an LLM via API), so the retrieved evidence supports the reasoning process and reduces hallucinations.\"}]","RRAML: Reinforced Retrieval Augmented Machine Learning | PDF",1785720221,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rraml-reinforced-retrieval-augmented-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rraml-reinforced-retrieval-augmented-machine-learning/118782/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does RRAML address in retrieval-augmented LLM usage?","Question",{"text":75,"@type":76},"RRAML targets the misalignment between retriever and reasoner, which can lead to irrelevant or even dangerous retrieved content and consequently hallucinations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RRAML reduce reliance on retraining or gradient access to LLMs?",{"text":80,"@type":76},"RRAML uses reinforcement learning to improve the system without needing access to LLM gradients, and it alleviates the need to retrain LLMs for specific tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the core design of RRAML?",{"text":84,"@type":76},"RRAML couples an efficient retriever that searches a large user-provided database with a reasoner (e.g., an LLM via API), so the retrieved evidence supports the reasoning process and reduces hallucinations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]