[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84832-en":3,"doc-seo-84832-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84832,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models","Large Language Models increasingly rely on external tools to complete complex tasks, yet they often struggle to manage large tool collections efficiently. Retrieval-based preselection can mitigate input-length and latency limits, but many retrievers are misaligned with tool-calling LLMs because they are trained separately. PORTS introduces an odds-ratio preference optimization strategy that fine-tunes a retriever using a perplexity-inspired preference signal from a frozen LLM and a contrastive semantic loss on documentation strings.","PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models  \nLorenzo Molfetta Giacomo Frisoni Nicolò Monaldini Gianluca Moro  \nDepartment of Computer Science and Engineering, University of Bologna {lorenzo.molfetta, giacomo.frisoni, [gianluca.moro}@unibo.it](gianluca.moro}@unibo.it), [nicolo.monaldini@studio.unibo.it](nicolo.monaldini@studio.unibo.it)  \narXiv :2607 .0544 1v 1 [ cs .IR] 3 Jul 2026  \nAbstract  \nIntegrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks.  \nSince LLMs still struggle to effectively manage large tool collections, researchers have begun exploring retrieval-based methods to preselect the most relevant options, addressing input length and latency constraints. However, existing retrievers are often misaligned with toolcalling LLMs due to their separate training processes. This paper presents PORTS, a novel odds ratio preference optimization method for training retrievers aimed at tool selection. Using a perplexity-inspired preference signal from a frozen LLM, our approach fine-tunes a retriever to find helpful tools by optimizing the correlation between the selection probabilities and the downstream performances while jointly enforcing a contrastive semantic loss between documentation strings. The versatility of PORTS and its ability to significantly improve tool selection accuracy are demonstrated through extensive experiments on six datasets, two encoder models, and three LLMs with diverse prior knowledge. With low computational demands, our alignment process facilitates generalization to new queries and tools, proving valuable for practical applications with evolving toolsets.1  \n1 Introduction  \n“The right tool for the right job.”—Proverb  \nEquipping Large Language Models (LLMs) with the capability to dynamically interact with external  \n1 Code, models, and datasets are publicly available at  \n[https://github.com/disi-unibo-nlp/ports](https://github.com/disi-unibo-nlp/ports)  \n*The definitive, copyrighted, peer-reviewed, and edited version of this article is published in the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 10007–10030, Suzhou, China, 2025 . Association for Computational Linguistics. [https://aclanthology.org/2025.emnlp-main](https://aclanthology.org/2025.emnlp-main). 507/ . DOI: 10.18653/v1/2025.emnlp-main.507 .  \nRoBERTa  \nBGE  \nTool Selection  \n Baseline  REPLUG  \n PORTS (Ours)  \n\n|  |  |\n| --- | --- |\n\n0 10 20 30 40 50 60 70 80  \nAvg. Recall (%)  \nFigure 1: Results overview. Comparison between frozen, REPLUG-tuned, and PORTS-tuned retrievers. Scores are averaged across Recall@{1,2,3} for three LLMs (if trained) and six datasets (test set) .  \ntools2 has garnered significant research attention. This integration not only improves the problemsolving potential of LLMs, but also dramatically expands their functional scope (Yao et al., 2022 ; Lazaridou et al., 2022) . When presented with a user query, tool-augmented LLMs can determine when and how to utilize specific tools to generate more accurate and informative responses. For example, tools can enable LLMs to use a calculator, set calendar events, and access real-time weather information. As the field continues to evolve, LLMs with tools are expected to play a pivotal role in shaping the future of Natural Language Processing (NLP) (Qu et al., 2024b) .  \nFine-tuning LLMs with tool usage examples is expensive and confines the acquired knowledge toa predefined set of tools (Qiao et al., 2023 ; Yang et al., 2023) . The in-context learning paradigm alleviates these issues, but the limitations in input length and noise for lengthy prompts make it  \n2Consistent with Qu et al. (2024b), we argue that all external means of augmenting LLMs should be classified as tools. Accordingly, we regard individual APIs as separate tools.  \nimpractical to manage many descriptions or demonstrations directly (Liu et al.,","cbCaiuUC7ibDDX8b","https://ap.wps.com/l/cbCaiuUC7ibDDX8b","pdf",1087022,5,1,24,"English","en",105,"# Introduction\n## Motivation and related work\n## Challenge: retriever–LLM misalignment\n## Approach overview","[{\"question\":\"What problem does PORTS address in tool-augmented LLM systems?\",\"answer\":\"PORTS targets the misalignment between retrievers and tool-calling LLMs, especially when large tool collections make it hard for the model to identify the most relevant tools efficiently.\"},{\"question\":\"How does PORTS train its retriever for tool selection?\",\"answer\":\"PORTS fine-tunes a retriever using a perplexity-inspired preference signal from a frozen LLM, optimizing correlation between selection probabilities and downstream performance while applying a contrastive semantic loss over tool documentation strings.\"},{\"question\":\"What evidence shows PORTS improves tool selection accuracy?\",\"answer\":\"The paper reports extensive experiments across six datasets, two encoder models, and three LLMs with diverse prior knowledge, demonstrating significant accuracy improvements while keeping computational demands low.\"}]",1784198604,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"ports-preference-optimized-retrievers-for-tool-selection-with-large-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ports-preference-optimized-retrievers-for-tool-selection-with-large-language-models/84832/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does PORTS address in tool-augmented LLM systems?","Question",{"text":76,"@type":77},"PORTS targets the misalignment between retrievers and tool-calling LLMs, especially when large tool collections make it hard for the model to identify the most relevant tools efficiently.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does PORTS train its retriever for tool selection?",{"text":81,"@type":77},"PORTS fine-tunes a retriever using a perplexity-inspired preference signal from a frozen LLM, optimizing correlation between selection probabilities and downstream performance while applying a contrastive semantic loss over tool documentation strings.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence shows PORTS improves tool selection accuracy?",{"text":85,"@type":77},"The paper reports extensive experiments across six datasets, two encoder models, and three LLMs with diverse prior knowledge, demonstrating significant accuracy improvements while keeping computational demands low.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]