[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-210528-en":3,"doc-seo-210528-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},210528,962084925502,"Emma Mercer","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","MARCA - A Checklist-Based Benchmark for Multilingual Web Search","Large language models increasingly act as information sources, but reliability hinges on effective web search, evidence selection, and omission-free answer synthesis. Existing benchmarks often emphasize English and tool use, while multilingual evaluation—especially Portuguese—remains limited. MARCA provides a bilingual English-Portuguese benchmark with 52 manually authored multi-entity questions and checklist-style rubrics measuring completeness and correctness. Fourteen models are tested in Basic and Orchestrator interaction settings.","arXiv :2604 . 14448v1 [ cs .CL] 15 Apr 2026  \nMARCA: A Checklist-Based Benchmark for Multilingual Web  \nSearch.  \nThales Sales Almeida 1 , Giovana Kerche Bonás 1 , Ramon Pires 1 , Celio Larcher2 , Hugo Abonizio 1 , Marcos Piau2 , Roseval Malaquias Junior 1 , Rodrigo Nogueira 1 , and Thiago  \nLaitz 1  \n1 Maritaca AI  \n2 Jusbrasil  \nAbstract  \nLarge language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesize complete answers. While recent benchmarks evaluate web-browsing and agentic tool use, multilingual settings, and Portuguese in particular, remain underexplored. We present MARCA, a bilingual (English and Portuguese) benchmark for evaluating LLMs on web-based information seeking. MARCA consists of 52 manually authored multi-entity questions, paired with manually validated checklist-style rubrics that explicitly measure answer completeness and correctness. We evaluate 14 models under two interaction settings: a Basic framework with direct web search and scraping, and an Orchestrator framework that enables task decomposition via delegated subagents. To capture stochasticity, each question is executed multiple times and performance is reported with run-level uncertainty. Across models, we observe large performance differences, find that orchestration often improves coverage, and identify substantial variability in how models transfer from English to Portuguese. The benchmark is available at [https://github.com/maritaca-ai/MARCA](https://github.com/maritaca-ai/MARCA)  \n1 Introduction  \nLarge language models (LLMs) are increasingly used as general-purpose sources of information: users ask for facts, comparisons, and up-to-date summaries, often expecting the system to both retrieve evidence and synthesize a reliable answer. This paradigm depends critically on models’ability to search and read the web effectively and to consolidate information without omissions or hallucinations.  \nRecent work has begun to evaluate such capabilities via web-browsing and agentic benchmarks [17, 28, 33, 8], but most evaluations remain centered on English. Portuguese, despite being one of the most widely spoken languages globally with over 250 million native speakers, has very limited benchmarks and studies that evaluate web search and evidence-grounded answering capabilities.  \nWe introduce MARCA (Maritaca AI Research Checklist evAluation), a benchmark for evaluating LLMs’ ability to find and verify information on the web in a multilingual setting. MARCA focuses on multi-entity questions that require systems to gather evidence from multiple webpages and produce structured, complete answers. Each question is paired with a manually curated checklist rubric, enabling fine-grained evaluation of correctness and coverage.  \nBeyond language, we study how interaction design shapes performance. We evaluate models in two frameworks: (i) a Basic setting where the model directly uses web_search and web_scrape, and (ii) an Orchestrator setting where the model can delegate sub-questions to subagents that perform web interactions.  \nOur contributions are:  \n• A new bilingual benchmark: parallel English and Portuguese versions of MARCA with 52 manually authored questions spanning 9 domains for evaluating web-search and evidencegrounded answering.  \n• Checklist-based evaluation: manually authored rubrics that measure both correctness and completeness on questions involving multiple entities.  \n• Framework comparison: an evaluation of Basic tool use versus Orchestrator-style delegation.  \n2 Related Work  \n2.1 Web Browsing and Search Benchmarks  \nEvaluating LLMs as systems that actively search and synthesize information from the web has received growing attention. WebGPT [17] introduced a browser-assisted QA paradigm where models issue queries, navigate pages, and produce grounded answers with human feedback. More recently, BrowseComp [28] proposed a benc","cbCaiqZGUJGDWqrO","https://ap.wps.com/l/cbCaiqZGUJGDWqrO","pdf",609641,1,14,"English","en",105,"# Introduction\n## Contributions\n# Related Work\n## Web Browsing and Search Benchmarks\n## Regional and Linguistic Variation in LLM Performance","[{\"question\":\"What problem does MARCA address for multilingual LLM information seeking?\",\"answer\":\"MARCA targets the gap in reliable evaluation for multilingual web-based question answering, especially for Portuguese, where prior benchmarks are limited or English-centric.\"},{\"question\":\"How is MARCA structured to evaluate LLM answers?\",\"answer\":\"MARCA uses 52 manually authored multi-entity questions paired with manually validated checklist-style rubrics that explicitly score both completeness and correctness.\"},{\"question\":\"What interaction frameworks are compared in the evaluation?\",\"answer\":\"Models are evaluated under two settings: a Basic framework using direct web search and scraping, and an Orchestrator framework that decomposes tasks via delegated subagents.\"}]","MARCA - A Checklist-Based Benchmark for Multilingual Web Search | PDF",1788665095,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"marca-a-checklist-based-benchmark-for-multilingual-web-search","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/marca-a-checklist-based-benchmark-for-multilingual-web-search/210528/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does MARCA address for multilingual LLM information seeking?","Question",{"text":76,"@type":77},"MARCA targets the gap in reliable evaluation for multilingual web-based question answering, especially for Portuguese, where prior benchmarks are limited or English-centric.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is MARCA structured to evaluate LLM answers?",{"text":81,"@type":77},"MARCA uses 52 manually authored multi-entity questions paired with manually validated checklist-style rubrics that explicitly score both completeness and correctness.",{"name":83,"@type":74,"acceptedAnswer":84},"What interaction frameworks are compared in the evaluation?",{"text":85,"@type":77},"Models are evaluated under two settings: a Basic framework using direct web search and scraping, and an Orchestrator framework that decomposes tasks via delegated subagents.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]