[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119970-en":3,"doc-seo-119970-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119970,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","RRAML - Reinforced Retrieval Augmented Machine Learning","Large language models (LLMs) have transformed machine learning by enabling strong language understanding and generation, yet typical API-based prompt usage limits context control and external source access. LLMs also produce hallucinations, motivating retrieval-augmented methods that use embeddings and vector databases. Current pipelines remain constrained by the need to manage embedding models and, when optimizing them, by access to LLM internals. RRAML integrates LLM reasoning with retriever-selected evidence, using reinforcement learning to avoid LLM gradients, reduce task-specific retraining, and align retrieval with generation while lowering hallucinations and harmful documents.","RRAML: Reinforced Retrieval Augmented Machine Learning  \nAndrea Bacciu1 , Florin Cuconasu1 , Federico Siciliano 1 , Fabrizio Silvestri 1 , Nicola Tonellotto2 and Giovanni Trappolini1,∗  \n1 Sapienza University of Rome 2 University of Pisa  \nAbstract  \nThe emergence of large language models (LLMs) has revolutionized machine learning and related fields, showcasing remarkable abilities in comprehending, generating, and manipulating human language. However, their conventional usage through API-based text prompt submissions imposes certain limitations in terms of context constraints and external source availability. LLMs suffer from the problem of hallucinating text, and in the last year, several approaches have been devised to overcome this issue: adding an external Knowledge Base or an external memory consisting of embeddings stored and retrieved by vector databases. In all the current approaches, though, the main issues are: (i) they need to access an embedding model and then adapt it to the task they have to solve; (ii) in case they have to optimize the embedding model, they need to have access to the parameters ofthe LLM, which in many cases are“black boxes”. To address these challenges, we propose a novel framework called Reinforced Retrieval Augmented Machine Learning (RRAML) . RRAML integrates the reasoning capabilities of LLMs with supporting information retrieved by a purpose-built retriever from a vast user-provided database. By leveraging recent advancements in reinforcement learning, our method effectively addresses several critical challenges. Firstly, it circumvents the need for accessing LLM gradients. Secondly, our method alleviates the burden of retraining LLMs for specific tasks, as it is often impractical or impossible due to restricted access to the model and the computational intensity involved. Additionally, we seamlessly link the retriever’s task with the reasoner, mitigating hallucinations and reducing irrelevant and potentially damaging retrieved documents. We believe that the research agenda outlined in this paper has the potential to profoundly impact the field of AI, democratizing access to and utilization of LLMs for a wide range of entities.  \nKeywords  \nDeep Learning, Information Retrieval, Large Language Models  \n1. Introduction  \nThe advent of Large Language Models (LLMs) has brought about a paradigm shift in machine learning and its related disciplines. LLMs [1, 2, 3, 4, 5] have exhibited unprecedented capabilities in understanding, generating, and manipulating the human language. Famously, ChatGPT [3] has entered the public space by reaching one million users in a matter of days. The way these models are used is through API that only allows submitting a textual prompt and getting back from the server the generated text. However, this causes an immediate limitation: all  \n22nd International Conference of the Italian Association for Artificial Intelligence (AIxIA 2023) -Discussion Papers ∗Corresponding author.  \n[Envelope-Open](Envelope-Open trappolini@diag.uniroma1.it)[ trappolini@diag.uniroma1.it](Envelope-Open trappolini@diag.uniroma1.it) (G. Trappolini)  \n © 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0) . CWPEURorkroceshopedings http://ceurISSN 1613-ws-0073.org CEUR Workshop Proceedings ([CEUR-WS.org](CEUR-WS.org))  \ninformation must be passed through this context, and we know transformer-based models do not scale nicely. Even if they did, API costs are charged on the basis of their usage. Therefore, using long contexts would be expensive. Even if one had the resources to run their own LLM, the costs of training and of the hardware infrastructure, and the environmental impact should be considered. There is an impendent need, though, to accommodate the enormous power of those models to specific user needs by making sure that they could use the reasoning capabilities of LLMs, through in-context learning [1] on t","cbCaiizAMz5saZoQ","https://ap.wps.com/l/cbCaiizAMz5saZoQ","pdf",295766,1,9,"English","en",105,"# 1. Introduction\n## Background on LLM capabilities and API limitations\n## Hallucinations and retrieval-augmented approaches\n## Misalignment between retriever and reasoner\n## Proposed framework: Reinforced Retrieval Augmented Machine Learning (RRAML)","[{\"question\":\"What problem does RRAML address in retrieval-augmented LLM usage?\",\"answer\":\"RRAML addresses hallucinations and the misalignment between the retriever and the reasoner, where retrieved information may be irrelevant or even dangerous to the final answer.\"},{\"question\":\"How does RRAML differ from prior retrieval-augmented methods?\",\"answer\":\"RRAML links the retriever task with the reasoner by integrating reinforcement-learning-based reasoning with supporting information retrieved from a user-provided database.\"},{\"question\":\"Why does RRAML reduce the need for fine-tuning?\",\"answer\":\"RRAML is designed to avoid relying on access to LLM gradients and to mitigate the need for retraining LLMs for specific tasks, which is often impractical or impossible due to restricted model access and compute costs.\"}]","RRAML - Reinforced Retrieval Augmented Machine Learning | PDF",1785727317,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"rraml-reinforced-retrieval-augmented-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/rraml-reinforced-retrieval-augmented-machine-learning/119970/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does RRAML address in retrieval-augmented LLM usage?","Question",{"text":76,"@type":77},"RRAML addresses hallucinations and the misalignment between the retriever and the reasoner, where retrieved information may be irrelevant or even dangerous to the final answer.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does RRAML differ from prior retrieval-augmented methods?",{"text":81,"@type":77},"RRAML links the retriever task with the reasoner by integrating reinforcement-learning-based reasoning with supporting information retrieved from a user-provided database.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does RRAML reduce the need for fine-tuning?",{"text":85,"@type":77},"RRAML is designed to avoid relying on access to LLM gradients and to mitigate the need for retraining LLMs for specific tasks, which is often impractical or impossible due to restricted model access and compute costs.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]