[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-186812-en":3,"doc-seo-186812-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},186812,1374402739827,"Nguyễn Văn Học","https://ap-avatar.wpscdn.com/avatar/14000c97e7351f1a627?x-image-process=image/resize,m_fixed,w_180,h_180&k=1787885694763230660",8,"Research & Report","2023 EACL Main - Benchmark Results for Question Answering Datasets","Benchmark results compare multiple question answering models across several datasets, reporting retrieval metrics such as R-1, R-2, R-L, and aggregated R-Lsum. Training, validation, and test splits are listed for WikiCQA, GooAQ-S, and ELI5, showing differing dataset scales. Model variants including QGb, QGours, and enhanced combinations are evaluated with fluency and correctness-related measures, and separate tables summarize performance for different prompt or answer types (e.g., QAb vs QAs).","| Datasets | Train | Val | Test |\n| --- | --- | --- | --- |\n| WikiCQA | 16,162 | 2,020 | 2,020 |\n| GooAQ-S | 200,000 | 500 | 498 |\n| ELI5 | 272,634 | 9,812 | 24,512 |\n\n\n| Title: How to Prepare a Healthy Meal for Your Pet Dog. Summary: To prepare a healthy meal for your dog, choose lean meat with the bones and fat removed, like chicken or beef   |\n| --- |\n| Q: What can I feed my dog if I have run out of dog food?\u003Cbr>A: In the short term, any bland human food such as chicken or white fish with rice or pasta is just fine . . .\u003Cbr>Q: How much homemade dog food do you feed your dog?\u003Cbr>A: Great question because it highlights one of the problems of feeding home prepared foods . . .\u003Cbr>Q: What should I not feed my dog?\u003Cbr>A: There are many human foods that are toxic to dogs. Top of the list of foods NOT to give are . . . |\n\n\n| Dataset Model |  | R-1 | R-2 | R-L | R-Lsum |\n| --- | --- | --- | --- | --- | --- |\n| WikiCQA | QGb QGours | 48.39 | 26.84 | 46.08 | 46.16 |\n|  |  | 49.22 | 27.79 | 46.98 | 47.08 |\n| GooAQ-S | QGb QGours | 44.26 | 19.73 | 41.11 | 41.08 |\n|  |  | 45.26 | 20.69 | 42.20 | 42.06 |\n| ELI5 | QGb QGours | 28.62 | 10.10 | 25.93 | 26.23 |\n|  |  | 29.15 | 10.36 | 26.40 | 26.69 |\n\n\n| Model | fluency\u003Cbr>score | % | relevance |  | correctness\u003Cbr>score % |\n| --- | --- | --- | --- | --- | --- |\n|  |  |  | score | % |  |\n| QGb | 4.83 | 8.3 | 3.84 | 19.0 | 3.49 19.3 |\n| QGours | 4.83 | 7.3 | 3.89 | 24.7 | 3.58 23.3 |\n\n\n| Dataset | Model | R-1 | R-2 | R-L | R-Lsum |\n| --- | --- | --- | --- | --- | --- |\n| WikiCQA | QAb | 18.41 | 5.53 | 14.72 | 15.98 |\n|  | QAs | 24.10 | 7.07 | 18.02 | 20.31 |\n| GooAQ-S | QAb | 18.58 | 5.80 | 14.77 | 15.66 |\n|  | QAs | 21.93 | 6.12 | 16.62 | 18.35 |\n| ELI5 | QAb | 12.38 | 2.19 | 8.98 | 10.63 |\n|  | QAs | 13.97 | 2.29 | 10.00 | 11.77 |\n\n\n| Dataset | Model | R-1 | R-2 | R-L | R-Lsum |\n| --- | --- | --- | --- | --- | --- |\n| WikiCQA | QAb+f | 27.16 | 7.32 | 19.85 | 23.49 |\n|  | QAs+f | 28.44 | 8.14 | 20.56 | 24.77 |\n| GooAQ-S | QAb+f | 27.72 | 7.68 | 20.08 | 23.64 |\n|  | QAs+f | 28.44 | 7.73 | 20.17 | 24.12 |\n| ELI5 | QAb+f | 22.06 | 3.93 | 13.93 | 19.72 |\n|  | QAs+f | 23.37 | 4.27 | 14.53 | 20.88 |\n\n\n| Models | R-1 | R-2 | R-L | R-LSum |\n| --- | --- | --- | --- | --- |\n| QGb | 48.39 | 26.84 | 46.08 | 46.16 |\n| QGb +CLs | 48.67 | 27.34 | 46.43 | 46.46 |\n| QGb +CLt | 48.72 | 27.26 | 46.33 | 46.41 |\n| QGb +CLs +AR | 49.26 | 27.77 | 46.82 | 46.88 |\n| QGb +CLt +AR | 49.22 | 27.79 | 46.98 | 47.08 |\n\n\n| A The most recent statistics show that around 2 .5% of small businesses are audited by the irs.\u003Cbr>QB Are small businesses audited by the irs?\u003Cbr>QG How many small businesses are audited by theirs?\u003Cbr>QR What percentage of small businesses are audited? |\n| --- |\n| A Try using poster putty to secure the images in your collage. if that doesn’t work, you may have to use nails.\u003Cbr>QB How do i put pictures in a collage?\u003Cbr>QG How do i make a collage without nails?\u003Cbr>QR How can i make a collage on a wall with a textured surface? |\n| (a) Good examples. |\n| A Foods such as popsicles, hard candy, and gelatin can be eaten on a clear liquid diet.\u003Cbr>QB what can you eat on a clear liquid diet?\u003Cbr>QG What can I eat to lose weight?\u003Cbr>QR What food can be included on a clear liquid diet? |\n| A You’ll have to remove the door and sand, prime, and repaint it.\u003Cbr>QB What do i have to do to make my bedroom look nice?\u003Cbr>QG How do i fix a broken door?\u003Cbr>QR What do i do if my door is sticking to the weather strip? |","cbCaicE42ZiXiUqp","https://ap.wps.com/l/cbCaicE42ZiXiUqp","pdf",416035,2,1,13,"English","en",105,"# Datasets\n## WikiCQA\n## GooAQ-S\n## ELI5\n# Model Evaluation Metrics\n## Fluency and Relevance\n## Correctness Score\n# Results by Dataset\n## R-1 R-2 R-L R-Lsum","[{\"question\":\"Which datasets are used for the question answering benchmark and how are they split?\",\"answer\":\"The document reports WikiCQA, GooAQ-S, and ELI5 with Train, Val, and Test counts. Each dataset uses its own scale across these splits.\"},{\"question\":\"What evaluation metrics are reported for model performance?\",\"answer\":\"Results include R-1, R-2, R-L, and R-Lsum. Additional tables also report fluency-related and correctness-related scores.\"},{\"question\":\"Which model variants are compared in the final results table?\",\"answer\":\"The document compares QGb and several enhanced variants such as QGb +CLs, QGb +CLt, and versions adding AR. These variants are listed with their corresponding R-1, R-2, R-L, and R-Lsum scores.\"}]","2023 EACL Main - Benchmark Results for Question Answering Datasets | PDF",1788377109,33,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"2023-eacl-main-benchmark-results-for-question-answering-datasets","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/2023-eacl-main-benchmark-results-for-question-answering-datasets/186812/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-05","2026-09-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Which datasets are used for the question answering benchmark and how are they split?","Question",{"text":76,"@type":77},"The document reports WikiCQA, GooAQ-S, and ELI5 with Train, Val, and Test counts. Each dataset uses its own scale across these splits.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What evaluation metrics are reported for model performance?",{"text":81,"@type":77},"Results include R-1, R-2, R-L, and R-Lsum. Additional tables also report fluency-related and correctness-related scores.",{"name":83,"@type":74,"acceptedAnswer":84},"Which model variants are compared in the final results table?",{"text":85,"@type":77},"The document compares QGb and several enhanced variants such as QGb +CLs, QGb +CLt, and versions adding AR. These variants are listed with their corresponding R-1, R-2, R-L, and R-Lsum scores.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]