[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160523-en":3,"doc-seo-160523-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},160523,2336474459895,"Gloria","https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916",8,"Research & Report","Learning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM","Calibration systems for high-stakes computerized adaptive tests (CATs) are essential for maintaining an operational item bank, yet new items enter with limited response data, making item parameters and resulting scores unreliable estimators of examinees’ true ability. This work introduces a pre-CAT approach that trains a neural network to generate low-dimensional item content embeddings, which then support an interpretable linear explanatory IRT model. A neural parameterization of the 3PL model learns discrimination and difficulty via embeddings, while the chance parameter is fixed; both item and ability are fit jointly using Monte Carlo EM, evaluated on Duolingo English Test tasks.","arXiv :2607 .06905v 1 [ stat .AP] 8 Jul 2026  \nLearning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM  \nJames Sharpnack, Kai-Ling Lo  \nDuolingo  \nJuly 9, 2026  \nAbstract  \nCalibration systems for high-stakes computerized adaptive tests (CATs) are essential for growing and maintaining the test’s operational item bank. When a new test item is added to an item bank, we have few response data to provide accurate item parameter estimates, leading to test scores that are poor estimators of the test taker’s true score. Item features and explanatory item response theory (IRT) models mitigate this impact by incorporating item information into the calibration process. Neural IRT models—IRT models where the item parameters are the output of a neural net — provide a powerful framework for accurate calibration, but learning hyperparameters and selecting neural architectures in real time while the CAT is scoring real test takers is impractical and a threat to score validity. In this work, we propose an initial step prior to launching a CAT that fits a neural net to produce low dimensional embeddings. The production calibration system can leverage these embeddings to use a simple and interpretable linear explanatory IRT model. We use a neural parameterization of the 3-parameter logistic (3PL) IRT model in which a feature network maps each item’s content features to a low-dimensional representation hj = z (xj) ∈ Rd , from which we model the discrimination and difficulty parameters (a, b) using generalized linear forms. Due to known identifiability issues when jointly estimating the chance parameter cand test taker ability [Lord, 1980], we set that to a global constant as opposed to making it also an output of the neural net. The feature network and the latent test-taker abilities θ are fit jointly via Monte Carlo Expectation-Maximization (MCEM), removing the need for a separate ability-estimation or pre-calibration stage. Using an item-split protocol that holds out entire items to simulate feature-only evaluation, we apply this approach to two task types from the Duolingo English Test practice test—yes/no vocabulary (Y/N Vocab) and vocabulary-in-context (ViC)—and search over feature sets, network architectures, and representation dimensions d. We find that a shallow two-layer ReLU network with d = 6 and hand-engineered scalar input features matches or outperforms larger architectures on held-out items for both task types. This work is a first step toward a compact, contentderived item embedding for use as a feature in the Scalable Parametric Item Calibration Engine (SPICE) [Nydick et al., 2026], the fully Bayesian calibration engine at the core of the S2A3 system for adaptive testing [Sharpnack et al., 2026] .  \n1 Introduction  \nItem response theory (IRT) models a test taker’s ability and item characteristics, known as item parameters, via parametric models of a test taker’s response to a given item [Lord, 1980, Wright and Stone, 1979] . A key advantage of IRT is interpretability: item parameters can be inspected directly to curate an item bank (e.g., ensuring a wide range of difficulties) and to construct computerized adaptive test (CAT) administration rules. However, traditional IRT calibration requires many responses per item—often hundreds—to meet the standards of a high-stakes test, typically collected during a piloting phase before an item is used for scoring. Piloting has costs: test takers spend time on items that do not count toward their score; items administered  \noutside the high-stakes test (e.g., on practice tests) are exposed to the public and pose a security risk [LaFlair et al., 2022, Way, 1998]; and test takers are less motivated to answer items to the best of their ability on unscored items [Cheng et al., 2014] . One way to calibrate new items with few responses is to extend IRT with item features derived from item content, such as NLP features or LLM embeddings [Fischer, 1973, McCarthy et al., ","cbCaivWJFTYQ5Qwy","https://ap.wps.com/l/cbCaivWJFTYQ5Qwy","pdf",576913,3,1,13,"English","en",105,"# Introduction\n## Feature-only calibration setting\n## Related work and computational psychometrics\n## Proposed approach and evaluation setup","[{\"question\":\"Why is CAT item calibration difficult when new items are added to the item bank?\",\"answer\":\"New items lack sufficient response data, making accurate item parameter estimates hard and causing CAT scores to be poor estimators of test takers’ true ability.\"},{\"question\":\"How does the proposed method use neural embeddings in IRT calibration?\",\"answer\":\"A feature network maps each item’s content features to a low-dimensional embedding, then discrimination and difficulty parameters are modeled using generalized linear forms within a neural-parameterized 3PL IRT framework.\"},{\"question\":\"How are the model components trained in the paper?\",\"answer\":\"The feature network and latent test-taker abilities are fit jointly using Monte Carlo Expectation-Maximization (MCEM), avoiding a separate ability-estimation or pre-calibration stage.\"}]","Learning Item Embeddings and Hyperparameters for IRT Calibration via Monte Carlo EM | PDF",1788068448,33,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"learning-item-embeddings-and-hyperparameters-for-irt-calibration-via-monte-carlo-em","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/learning-item-embeddings-and-hyperparameters-for-irt-calibration-via-monte-carlo-em/160523/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-03","2026-08-30",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is CAT item calibration difficult when new items are added to the item bank?","Question",{"text":76,"@type":77},"New items lack sufficient response data, making accurate item parameter estimates hard and causing CAT scores to be poor estimators of test takers’ true ability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed method use neural embeddings in IRT calibration?",{"text":81,"@type":77},"A feature network maps each item’s content features to a low-dimensional embedding, then discrimination and difficulty parameters are modeled using generalized linear forms within a neural-parameterized 3PL IRT framework.",{"name":83,"@type":74,"acceptedAnswer":84},"How are the model components trained in the paper?",{"text":85,"@type":77},"The feature network and latent test-taker abilities are fit jointly using Monte Carlo Expectation-Maximization (MCEM), avoiding a separate ability-estimation or pre-calibration stage.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]