[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84186-en":3,"doc-seo-84186-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84186,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","From Text to Parameters Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings","Newly developed items require field testing to estimate psychometric properties, creating a cold-start obstacle for item calibration. This work evaluates predicting item parameters directly from item text embeddings using regularized regression and repeated cross-validated R2 with uncertainty. It introduces two performance upper bounds: a reliability ceiling from parameter standard errors and a design ceiling from simulation-based power calibration. Results on a mathematics bank and a medical-licensure benchmark show difficulty can be predicted with substantial reliability recovery, while discrimination and pseudo-guessing are limited, and ceiling-aware metrics guide benchmark design.","FROM TEXT TO PARAMETERS 1  \nFrom Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings  \nShi-Ting Chen 1 and Jinsong Chen 1  \n1 Faculty of Education, The University of Hong Kong  \narXiv :2607 .07 14 1v 1 [ cs .CL] 8 Jul 2026  \nAuthor Note  \nShi-Ting Chen and Jinsong Chen contributed equally to this work and share first authorship. Corresponding author: Jinsong Chen ([jinsong.chen@live.com](jinsong.chen@live.com)).  \nFROM TEXT TO PARAMETERS 2  \nAbstract  \nNewly developed items must ordinarily be field tested before their psychometric properties are known, creating a cold-start problem for item calibration. Predicting item parameters from item features is a long-standing measurement problem dating back to the linear logistic test model (LLTM); modern text embeddings automate the design matrix that the LLTM tradition specified by hand. We propose an evaluation framework that combines regularized regression on item-text embeddings, repeated cross-validated R 2 reported with its resampling standard deviation, and two upper bounds on attainable performance: a reliability ceiling derived from the standard errors of the estimated parameters, and a design ceiling derived from a simulation-based power calibration of the prediction pipeline. Applying the framework to a mathematics item bank (EEDI) and a medical-licensure difficulty benchmark (BEA 2024), we find that item difficulty is substantially predictable from text (repeated cross-validated R2 = 0 .53, about 57% of its reliability ceiling), whereas discrimination and pseudo-guessing appear progressively less predictable. Once read against the ceilings, however, this hierarchy is revealed to be largely a hierarchy of target reliability rather than of text signal: text recovers a nearly uniform 57–63% of the reliable variance in every difficulty target, while the 3PL pseudo-guessing parameter has a reliability ceiling of essentially zero and is simply an unusable target at this calibration precision. On BEA, embedding-based regression matches leaderboard-level RMSE while explaining almost no variance, illustrating why scale-free metrics and explicit ceilings are needed to benchmark this task. All results use repeated cross-validation on the pooled items; we show that a single train–test split can inflate apparent accuracy by 0.1–0.15 in R2 , and discuss implications for calibration-support applications and for how item-difficulty benchmarks should be constructed and reported.  \nKeywords: item difficulty modeling, text embeddings, regularized regression, linear logistic test model, item calibration, cross-validation  \nFROM TEXT TO PARAMETERS 3  \nFrom Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings  \nIntroduction  \nOperational assessment programs rely on the continuous development and replenishment of item pools. However, newly developed items face a cold-start problem. Traditionally, estimating item parameters requires administering new items to a sample of test takers, collecting their responses, and then fitting a psychometric model, such as an item response theory (IRT) model, to obtain parameter estimates (McCarthy et al., 2021) . This field testing process is costly, time-consuming, and may expose items before operational use (Yancey, Runge, Laflair, & Mulcaire, 2024) . The burden is most acute exactly where the demand for fresh items is greatest: high-stakes programs that must retire exposed items quickly, and computerized adaptive testing (CAT) systems whose large item banks cannot function until every item carries calibrated parameters. Accurate item parameter prediction from item content would allow new items to enter operational use sooner, reduce pretest sample requirements, and support item development itself, for example by flagging items whose predicted difficulty departs from the intended blueprint. The measurement community has begun to take this","cbCaiuFw1Ep3vIsa","https://ap.wps.com/l/cbCaiuFw1Ep3vIsa","pdf",386573,5,1,40,"English","en",105,"# From Text to Parameters: Predicting Item Parameters from Embedding Regularization with Reliability and Design Ceilings\n## Abstract\n## Introduction","[{\"question\":\"What do the experiments show about predicting different IRT parameters from text?\",\"answer\":\"Item difficulty is substantially predictable from text and can recover a large portion of the reliability ceiling. In contrast, discrimination and pseudo-guessing become progressively less predictable, with pseudo-guessing having an essentially zero reliability ceiling under the paper’s calibration precision.\"}]",1784193771,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"from-text-to-parameters-predicting-item-parameters-from-embedding-regularization-with-reliability-and-design-ceilings","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-text-to-parameters-predicting-item-parameters-from-embedding-regularization-with-reliability-and-design-ceilings/84186/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What do the experiments show about predicting different IRT parameters from text?","Question",{"text":76,"@type":77},"Item difficulty is substantially predictable from text and can recover a large portion of the reliability ceiling. In contrast, discrimination and pseudo-guessing become progressively less predictable, with pseudo-guessing having an essentially zero reliability ceiling under the paper’s calibration precision.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,102,107,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":22,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":20,"slug":129},19,"General","general"]