[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84885-en":3,"doc-seo-84885-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84885,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Measuring the Practice of Shared-Decision Making (OPTION12) An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability","Shared decision-making (SDM) in clinical consultations is commonly measured using the OPTION12 instrument, where 12 items are observer-rated on a 5-point Likert scale. Human coding is time-intensive and often suffers from coder disagreement. The study investigates whether open-source privacy-preserving smaller LLMs (OS-sLLMs) can perform OPTION12 scoring with humans in the loop. Using Dutch melanoma consultation transcripts, models are evaluated via prompt refinement, correlation analyses, and judge-LLM-based disagreement resolution.","Measuring the practice of shared-decision making (OPTION12): An Investigation into Open-sourced Smaller LLMs (OS-sLLMs) for Better Privacy and Sustainability  \nTamara Wit 1†, Lifeng Han 1 ,2†, Carly Heipon 1 , David Lindevelt2  \nAnne Stiggelbout 1 , Suzan Verberne2  \nOn behalf of the 4D PICTURE consortium  \nBDS, Leiden University Medical Centre, NL The Leiden Institute of Advanced Computer Science, LU, NL † co-first, corresponding: {t.wit, l. han} @ lumc. nl  \narXiv :2607 .06 127v2 [ cs .CL] 9 Jul 2026  \nAccepted Abstract  \nBackground: Shared decision-making (SDM) is important in clinical consultations in which patients discuss and decide on treatment options with clinicians. This SDM process is often coded with the OPTION12 instrument, an observer-based tool, consisting of 12 items, rated on a 5-point Likert scale. This coding is conducted by human coders. Coding is timeintensive and frequently accompanied by disagreement between coders. We explore the capability of open-source privacy-preserving smaller LLMs (OS-sLLMs) to perform the coding task, potentially automating the process, with humans in the loop. Methods: 26 transcripts of Dutch melanoma patients consultations with clinicians were double-coded by two coders. Two human coders resolved their disagreements after independent coding. To evaluate OPTION12 coding using OS-sLLMs, the consultation data was divided into development and testing sets. We designed the complete investigation framework (see Figure 1) . It includes 1) a pilot study of the development set (11 interviews) for prompt-refinement and sLLM selection. 2) deployment of the fine-tuned prompts and best performing sLLM (judge-sLLM) augmented with few-shot examples from the development set on the testing set (15 interviews) and asking the judgesLLM to resolve the disagreement on other OS-sLLMs’ scores. The rationale for judge-llm: it behaves like a third-party human to resolve the disagreements of annotators. Instead of using humans, a judge-LLM is consulted to resolve disagreement. This is designed to mimic how human annotators resolve disagreements, by discussing until consensus is achieved. We plan to investigate alternative model-merging methods for the future. First, for prompt re-  \nfinement, we use chain-of-thoughts (CoTs), LLM-assisted prompting, human-in-the-loop with sample output (few-shots) feedback. For OS-sLLMs, we use both 1) general domain models Llama, Gemma, and Mistral7b, and 2) medical domain models Meditron and Medllama. Second, for judge-sLLM selection, on the system-level, we measure the overall correlation of each sLLM and human coding using Spearman and Pearson correlation scores. On the segment-level, we also look into the most agreed-upon and disagreed-upon items . We will perform both qualitative and quantitative analysis by discussing the evaluation scoresand categorise the OS-sLLMs’ behaviours with examples. Finally, we deploy the OS-sLLMs on the testing set and use the judge-sLLM to resolve disagreements.  \nPreliminary results: Five OS-sLLMs show the following findings:  \n• Three general domain OS-sLLMs perform better than the two medical domain ones, which both generate hallucinations and do not follow prompts precisely, indicating further developments is needed for medical OS-sLLMs.  \n• Mistral7b outperformed the other two Gemma3:12b and Llama3 . 1:8b by 4 consensus with human coding, vs 3 items.  \n• The overall correlation with human coding on these 12 items is (0.83, 0.80, 0.64) using Pearson correlation, and (0.81, 0.78, 0 .61) using Spearman rank correlation from the three models (gemma3:12b, llama3.1:8b, mistral7b) .  \n• For the items for which OS-sLLMs agree with human coders, OS-sLLMs can generate the same sentences as humans in  \nsome cases, but in other cases, they generate better quotes than humans.  \nImpacts: This study reports the first research findings using OS-sLLMs on scoring SDM with the OPION12 in Dutch melanoma patients’ consultation transcripts. The perform","cbCaihvCdUEEgHhs","https://ap.wps.com/l/cbCaihvCdUEEgHhs","pdf",395586,1,14,"English","en",105,"# Background\n## SDM and OPTION12\n# Methods\n## Data and double coding\n## OS-sLLM deployment and judge-LLM\n# Preliminary Results\n## Model comparisons and correlations\n## Item-level agreement\n# Impacts\n## Automation potential and human-in-the-loop deployment","[{\"question\":\"What problem does the OPTION12 coding process face in practice?\",\"answer\":\"OPTION12 coding relies on human observers and is time-intensive. It also frequently involves disagreements between coders, motivating automated support.\"},{\"question\":\"How does the study use open-source smaller LLMs to score SDM?\",\"answer\":\"The work refines prompts and selects OS-sLLMs on a development set, then deploys the best-performing model with few-shot examples on a testing set. A judge-LLM resolves disagreements to mimic third-party human adjudication.\"},{\"question\":\"What do preliminary results indicate about which models perform best?\",\"answer\":\"Three general-domain OS-sLLMs outperform two medical-domain models that produce hallucinations and fail to follow prompts precisely. Mistral7b shows the strongest consensus with human coding among the compared models, with high overall correlation on the 12 items.\"}]",1784199020,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"measuring-the-practice-of-shared-decision-making-option12-an-investigation-into-open-sourced-smaller-llms-os-sllms-for-better-privacy-and-sustainability","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/measuring-the-practice-of-shared-decision-making-option12-an-investigation-into-open-sourced-smaller-llms-os-sllms-for-better-privacy-and-sustainability/84885/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the OPTION12 coding process face in practice?","Question",{"text":75,"@type":76},"OPTION12 coding relies on human observers and is time-intensive. It also frequently involves disagreements between coders, motivating automated support.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study use open-source smaller LLMs to score SDM?",{"text":80,"@type":76},"The work refines prompts and selects OS-sLLMs on a development set, then deploys the best-performing model with few-shot examples on a testing set. A judge-LLM resolves disagreements to mimic third-party human adjudication.",{"name":82,"@type":73,"acceptedAnswer":83},"What do preliminary results indicate about which models perform best?",{"text":84,"@type":76},"Three general-domain OS-sLLMs outperform two medical-domain models that produce hallucinations and fail to follow prompts precisely. Mistral7b shows the strongest consensus with human coding among the compared models, with high overall correlation on the 12 items.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]