[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-191390-105":3,"detail-sidebar-cat-1-en-105":80,"doc-detail-191390-en":126},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":73,"head_meta":75,"extra_data":77,"updated_unix":79},105,"en","conversation-analytics-for-product-improvement-model-performance-and-failure-analysis","Conversation Analytics for Product Improvement - Model Performance and Failure Analysis","","Conversation analytics framework evaluates how AI systems handle user emotions, intent, linguistic style, and trust, safety, and ethics. It measures general usage statistics across conversations and query pairs, including distribution by intent, feedback, language expression, ethics, and dialogue dynamics. Model quality is compared using ranking metrics (NDCG@10, R@10, P@10) across turn, sliding chunk, and session inference/ingestion time. Challenge typologies show role recognition, progression, and semantic-context failures with concrete query examples.",{"@graph":14,"@context":72},[15,34,55],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/template/","Template",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/template/presentations/","Presentations",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/template/conversation-analytics-for-product-improvement-model-performance-and-failure-analysis/191390/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/conversation-analytics-for-product-improvement-model-performance-and-failure-analysis/191390.png","ImageObject",442,249,{"name":42,"@type":43},"Valentina","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-29","2026-09-03",true,{"@type":52,"interactionType":53,"userInteractionCount":33},"InteractionCounter",{"@type":54},"ViewAction",{"@type":56,"mainEntity":57},"FAQPage",[58,64,68],{"name":59,"@type":60,"acceptedAnswer":61},"What dimensions are analyzed to understand conversation quality?","Question",{"text":62,"@type":63},"The analysis covers Emotion & Feedback, Intent & Purpose, Conversation Dynamics, Trust, Safety & Ethics, and Linguistic Style & Expression, linking each dimension to product-relevant insights.","Answer",{"name":65,"@type":60,"acceptedAnswer":66},"How is model performance measured in the document?",{"text":67,"@type":63},"Performance is reported with ranking metrics including NDCG@10, R@10, and P@10, across different settings such as turn, sliding chunk (k=3), session behavior, and corresponding inference and ingestion times.",{"name":69,"@type":60,"acceptedAnswer":70},"What are the main reasons models fail shown in the document?",{"text":71,"@type":63},"The document highlights role recognition failures, dynamic progression misses (static satisfaction vs gradual improvement), and semantic contextual misinterpretation where keyword matches ignore the user’s true service context.","https://schema.org",{"og:url":32,"og:type":74,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":76,"canonical":32},"index,follow",{"doc_id":78,"site_id":7},191390,1788408330,{"code":4,"msg":81,"data":82},"success",[83,87,92,97,102,107,112,117,122],{"id":84,"doc_module":22,"doc_module_name":25,"category_name":29,"show_sort_weight":85,"slug":86},11,90,"presentations",{"id":88,"doc_module":22,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},12,"Resumes",80,"resumes",{"id":93,"doc_module":22,"doc_module_name":25,"category_name":94,"show_sort_weight":95,"slug":96},14,"Invoices",70,"invoices",{"id":98,"doc_module":22,"doc_module_name":25,"category_name":99,"show_sort_weight":100,"slug":101},15,"Posters",60,"posters",{"id":103,"doc_module":22,"doc_module_name":25,"category_name":104,"show_sort_weight":105,"slug":106},16,"Social Media",50,"social-media",{"id":108,"doc_module":22,"doc_module_name":25,"category_name":109,"show_sort_weight":110,"slug":111},17,"Forms",40,"forms",{"id":113,"doc_module":22,"doc_module_name":25,"category_name":114,"show_sort_weight":115,"slug":116},18,"Letters",30,"letters",{"id":118,"doc_module":22,"doc_module_name":25,"category_name":119,"show_sort_weight":120,"slug":121},21,"Paper Templates",5,"papers-templates",{"id":123,"doc_module":22,"doc_module_name":25,"category_name":124,"show_sort_weight":4,"slug":125},158,"General","general-158",{"code":4,"msg":81,"data":127},{"doc_id":78,"user_id":128,"nickname":42,"user_avatar":129,"doc_module":22,"category_id":84,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":33,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":135,"language":136,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":12,"update_tm":79,"read_time":140},13056703020460,"https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923","| Analytical Area | Description | Product Insights |\n| --- | --- | --- |\n| Emotion & Feedback | Identifying users’ emotional states and feedback in conversations | Revealing satisfaction patterns and pain points for product improvement |\n| Intent & Purpose | Recognizing user intentions and goals | Evaluating alignment between intended and actual AI system usage |\n| Conversation Dynamics | Analyzing conversation flow, turn structure and resolution patterns | Identifying conversation bottlenecks and improving dialogue completion rates |\n| Trust, Safety & Ethics | Exploring trust-building and ethical issues in conversations | Identifying system reliability concerns and potential safety risks |\n| Linguistic Style & Expression | Analyzing language patterns and comprehension challenges | Helping calibrate system language to user comprehension levels |\n\n| General Statistics |  |\n| --- | --- |\n| Number of conversations\u003Cbr>Number of queries\u003Cbr>Avg. messages per conversation\u003Cbr>Avg. tokens per conversation Avg. relevant convs per query Total query-conversation pairs | 9,146\u003Cbr>1,583 5.4\u003Cbr>464\u003Cbr>20.44\u003Cbr>32,357 |\n| Query Task Distribution (%) |  |\n| Intent & Purpose\u003Cbr>Emotion & Feedback Linguistic Style & Expression Trust, Safety & Ethics Conversation Dynamics | 36.1%\u003Cbr>20.1%\u003Cbr>15.9%\u003Cbr>14.6%\u003Cbr>13.4% |\n\n\n| Model | Turn |  |  | Sliding chunk (k=3) |  |  | Session |  |  | Inference (s) | Ingestion (s) |\n| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |\n|  | NDCG@10 | R@10 | P@10 | NDCG@10 | R@10 | P@10 | NDCG@10 | R@10 | P@10 |  |  |\n| Commercial API Models |  |  |  |  |  |  |  |  |  |  |  |\n| Voyage-3-large | 0.5079 | 0.2609 | 0.4359 | 0.5063 | 0.2582 | 0.4327 | 0.5036 | 0.2615 | 0.4358 | 375.03 | 2620.00 |\n|  |  |  |  |  |  |  |  |  |  |  |  |\n| Text-embedding-3-large | 0.5078 | 0.2698 | 0.4389 | 0.5130 | 0.2696 | 0.4416 | 0.4876 | 0.2529 | 0.4190 | 245.90 | 1433.12 |\n| Text-embedding-3-small | 0.4897 | 0.2558 | 0.4183 | 0.4855 | 0.2558 | 0.4171 | 0.4664 | 0.2412 | 0.3972 | 196.99 | 1253.68 |\n| Embed-english-v3.0 | 0.4189 | 0.2237 | 0.3547 | 0.2547 | 0.1351 | 0.2116 | 0.3620 | 0.1923 | 0.2987 | 83.75 | 520.00 |\n| Open Source Models |  |  |  |  |  |  |  |  |  |  |  |\n| Stella_en_1.5B_v5 | 0.4907 | 0.2592 | 0.4141 | 0.4894 | 0.2528 | 0.4078 | 0.4722 | 0.2481 | 0.3961 | 7.84 | 336.92 |\n| Stella_en_400M_v5 | 0.4682 | 0.2490 | 0.3963 | 0.4651 | 0.2462 | 0.3919 | 0.4583 | 0.2400 | 0.3846 | 5.51 | 119.24 |\n| Jasper_en_vision_language_v1 | 0.4379 | 0.2317 | 0.3712 | 0.4309 | 0.2245 | 0.3615 | 0.4561 | 0.2382 | 0.3814 | 7.86 | 355.56 |\n| NV-Embed-v2 | 0.3170 | 0.2008 | 0.3251 | 0.3956 | 0.1988 | 0.3262 | 0.4592 | 0.2344 | 0.3855 | 13.92 | 279.73 |\n| NV-Embed-v1 | 0.2467 | 0.1226 | 0.1956 | 0.2603 | 0.1302 | 0.2080 | 0.4389 | 0.2242 | 0.3634 | 13.79 | 280.23 |\n| SFR-Embedding-2_R | 0.3344 | 0.1775 | 0.2805 | 0.3127 | 0.1639 | 0.2589 | 0.4474 | 0.2280 | 0.3722 | 10.40 | 213.57 |\n| Jina-embeddings-v3 | 0.3803 | 0.2053 | 0.3160 | 0.3983 | 0.2142 | 0.3363 | 0.3718 | 0.1995 | 0.3088 | 7.01 | 106.25 |\n| Modernbert-embed-base | 0.3594 | 0.1923 | 0.3026 | 0.3398 | 0.1795 | 0.2857 | 0.3579 | 0.1906 | 0.3016 | 7.49 | 45.63 |\n| Gte-Qwen2-1.5B-instruct | 0.4646 | 0.2412 | 0.3952 | 0.4386 | 0.2261 | 0.3708 | 0.3615 | 0.1919 | 0.2987 | 7.59 | 336.73 |\n| Gte-large-en-v1.5 | 0.3310 | 0.1821 | 0.2792 | 0.3246 | 0.1778 | 0.2726 | 0.3429 | 0.1840 | 0.2860 | 4.39 | 133.29 |\n| Bge-large-en-v1.5 | 0.3276 | 0.1757 | 0.2719 | 0.3105 | 0.1659 | 0.2539 | 0.3071 | 0.1617 | 0.2476 | 3.97 | 94.29 |\n| Cde-small-v2 | 0.1163 | 0.0606 | 0.0975 | 0.1226 | 0.0640 | 0.1007 | 0.0830 | 0.0463 | 0.0701 | 7.50 | 39.57 |\n\n\n| Challenge Type |  | Query Example | Incorrectly Retrieved Results | Why Models Fail |  |\n| --- | --- | --- | --- | --- | --- |\n| Role Recognition Failure | Assistant shares parenting and childcare advice user: Welcome to the parent teacher conference. So what is your child’s name?\u003Cbr>assistant: Megan Jones.\u003Cbr>user: She’s been having som","cbCailKxHv3C2a5S","https://ap.wps.com/l/cbCailKxHv3C2a5S","pdf",1890943,24,"English","# Analytical Areas\n## Emotion & Feedback\n## Intent & Purpose\n## Conversation Dynamics\n## Trust, Safety & Ethics\n## Linguistic Style & Expression\n\n# General Statistics\n## Query Task Distribution\n\n# Model Performance Comparison\n## Commercial API Models\n## Open Source Models\n\n# Challenge Types and Failure Modes\n## Role Recognition Failure\n## Dynamic Progression Failure\n## Semantic Contextual Misinterpretation","[{\"question\":\"What dimensions are analyzed to understand conversation quality?\",\"answer\":\"The analysis covers Emotion \\u0026 Feedback, Intent \\u0026 Purpose, Conversation Dynamics, Trust, Safety \\u0026 Ethics, and Linguistic Style \\u0026 Expression, linking each dimension to product-relevant insights.\"},{\"question\":\"How is model performance measured in the document?\",\"answer\":\"Performance is reported with ranking metrics including NDCG@10, R@10, and P@10, across different settings such as turn, sliding chunk (k=3), session behavior, and corresponding inference and ingestion times.\"},{\"question\":\"What are the main reasons models fail shown in the document?\",\"answer\":\"The document highlights role recognition failures, dynamic progression misses (static satisfaction vs gradual improvement), and semantic contextual misinterpretation where keyword matches ignore the user’s true service context.\"}]","Conversation Analytics for Product Improvement - Model Performance and Failure Analysis | PDF",8]