[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118857-en":3,"doc-seo-118857-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118857,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Testing machine learning systems in real estate - Applied software system testing for ML models","Uncertainty about how machine learning (ML) models produce estimates slows down the use of ML-enabled solutions in real estate. The central challenge is model transparency: how practitioners can ensure that ML systems deliver intended outputs and do not violate legal requirements. The article argues for a dedicated software testing framework for applied ML. It then shows how system-testing procedures can confirm that ML image classifiers used in automated valuation models perform as specified.","DOI: 10.1111/1540-6229.12416  \nORIGINAL ARTICLE  \nTesting machine learning systems in real estate  \nWayne Xinwei Wan1   Thies Lindenthal2   \n1 Department of Banking and Finance, Monash University, Clayton, Victoria, Australia  \n2 Department of Land Economy, The University of Cambridge, Cambridge, Cambridgeshire, UK  \nCorrespondence  \nWayne Xinwei Wan, Department of Banking and Finance, Monash University, W1025 Menzies Bldg, Wellington Rd, Clayton, VIC 3800, Australia. [Email: wayne.wan@monash.edu](Email: wayne.wan@monash.edu)  \nAbstract  \nUncertainty about the inner workings of machine learning (ML) models holds back the application of MLenabled systems in real estate markets. How do ML models arrive at their estimates? Given the lack of model transparency, how can practitioners guarantee that ML systems do not run afoul of the law? This article first advocates a dedicated software testing framework for applied ML systems, as commonly found in computer science. Second, it demonstrates how system testing can verify that applied ML models indeed perform as intended. Two system-testing procedures developed for ML image classifiers used in automated valuation models (AVMs) illustrate the approach.  \nKEYWORDS  \naccountability gap, computer vision, explainable machine learning, real estate, system testing  \n1  INTRODUCTION  \nThe black-box nature of machine learning (ML) techniques poses a risk to businesses developing ML-enabled systems: How can they verify that a system is performing in thewayit is meant to perform? Can they ensure that its outcomes are not spurious, biased, or unlawful? Still, ML-enabled systems continue to reshape commerce, personal interactions, entertainment, medicine, government services, state supervision—and research (Simester et al., 2020) . In real estate and urban studies, a rapidly expanding literature explores the potential of ML algorithms, introducing novel measurements of the physical environments or using these estimates to improve the traditional real estate valuation and urban planning processes (e.g., Glaeser et al., 2018; Johnson et al., 2020; Karimi et al., 2019; Lindenthal & Johnson, 2021; Liu et al., 2017; Rossetti et al., 2019; Schmidt &  \n© 2022 American Real Estate and Urban Economics Association.  \nLindenthal, 2020; Shen & Ross, 2021). These studies, again and again, demonstrate the undisputed power of ML systems as prediction machines. Still, it remains difficult for researchers to establish causality or for end users to understand the internal workings of any models. An “accountability gap”(Adadi & Berrada, 2018) remains: How do the models arrive at their prediction results? Can we trust them not to bend rules or to cut corners?  \nThis accountability gap holds back the deployment of ML-enabled systems in real-life situations (Ibrahim et al., 2020). If engineers cannot observe the inner workings ofthe models, how can they guarantee reliable outcomes? Furthermore, the accountability gap also leads to obvious dangers: Flaws in prediction machines are not easily discernible by classic cross-validation approaches (Ribeiro et al., 2016) . More importantly, the opacity of the ML models also gives rise to the legal and ethical concerns for its real-life applications (Mullainathan & Obermeyer, 2017). For instance, anecdotal evidence reports that some ML engines for recruitment have exerted biases again the female applicants (Dastin, 2018) . Traditional ML model validation metrics such as the magnitude of prediction errors or 􀀂1-scores can evaluate the models’predictive performance, but they provide limited insights for addressing the accountability gap.  \nTraining ML models is a software development process at heart. We therefore suggest that ML system developers follow best practices and industry standards in software testing. Particularly, the system-testing stage of software test regimes is essential: It verifies whether an integrated system performs the exact function as required in the init","cbCairUIRxJaoYhl","https://ap.wps.com/l/cbCairUIRxJaoYhl","pdf",1277203,1,25,"English","en",105,"# Introduction\n# Accountability gap and risks of black-box ML\n# Software testing framework for applied ML\n# System testing for ML image classifiers in AVMs\n## Local explanation and scaling challenges","[{\"question\":\"为什么房地产场景中应用ML模型会遇到“问责差距”？\",\"answer\":\"ML模型的黑箱特性使得难以追踪预测结果的形成过程，从而难以判断其是否遵守规则或法律要求。问责差距因此阻碍了ML在真实情境中的部署。\"},{\"question\":\"文章提出用什么方法提升应用ML系统的可信度？\",\"answer\":\"提出建立面向应用ML的专门软件测试框架，并强调系统测试阶段的重要性。系统测试用于验证集成系统是否按最初设计的要求执行，从而提升过程透明度与结果可信性。\"},{\"question\":\"文章如何说明系统测试能验证ML模型是否按预期工作？\",\"answer\":\"通过两类用于自动估值模型（AVM）的ML图像分类器用例，展示系统测试程序如何验证模型确实实现了预期功能。\"},{\"question\":\"为什么仅依赖模型可解释性方法不足以进行大规模验证？\",\"answer\":\"许多局部解释工具需要对每个观测进行人工检查，通常以定性方式呈现。由于缺乏可扩展性，这类工具难以在大样本场景下完成模型验证。\"}]","Testing machine learning systems in real estate - Applied software system testing for ML models | PDF",1785720646,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"testing-machine-learning-systems-in-real-estate-applied-software-system-testing-for-ml-models","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/testing-machine-learning-systems-in-real-estate-applied-software-system-testing-for-ml-models/118857/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"为什么房地产场景中应用ML模型会遇到“问责差距”？","Question",{"text":75,"@type":76},"ML模型的黑箱特性使得难以追踪预测结果的形成过程，从而难以判断其是否遵守规则或法律要求。问责差距因此阻碍了ML在真实情境中的部署。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文章提出用什么方法提升应用ML系统的可信度？",{"text":80,"@type":76},"提出建立面向应用ML的专门软件测试框架，并强调系统测试阶段的重要性。系统测试用于验证集成系统是否按最初设计的要求执行，从而提升过程透明度与结果可信性。",{"name":82,"@type":73,"acceptedAnswer":83},"文章如何说明系统测试能验证ML模型是否按预期工作？",{"text":84,"@type":76},"通过两类用于自动估值模型（AVM）的ML图像分类器用例，展示系统测试程序如何验证模型确实实现了预期功能。",{"name":86,"@type":73,"acceptedAnswer":87},"为什么仅依赖模型可解释性方法不足以进行大规模验证？",{"text":88,"@type":76},"许多局部解释工具需要对每个观测进行人工检查，通常以定性方式呈现。由于缺乏可扩展性，这类工具难以在大样本场景下完成模型验证。","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]