[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84678-en":3,"doc-seo-84678-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84678,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Addressing Predicate Redundancy in Research Knowledge Graphs","Research Knowledge Graphs (RKGs) support structured representation of scientific knowledge, yet weakly enforced schemas make them vulnerable to inconsistencies—especially around predicates. Duplicate predicates, represented by different identifiers for the same or highly similar relationships, create semantic redundancy, reduce reuse, and lower overall RKG quality. This paper introduces a framework that detects, resolves, and prevents duplicate predicates using similarity-based automation plus human validation. Implemented in ORKG, it extends SciKGDash with embedding clustering and resolution actions, and evaluates redundancy coverage and root modeling patterns.","arXiv :2607 .03 197v2 [ cs .DL] 7 Jul 2026  \n1  \nAddressing Predicate Redundancy in Research Knowledge Graphs: Duplicate Detection, Resolution, and Prevention  \nLena John 1[0009−0007−2097−9761], Sushant Aggarwal2 , Sören Auer 1[0000−0002−0698−2864], and Oliver Karras 1[0000−0001−5336−6899]  \nTIB-Leibniz Information Centre for Science and Technology, Hannover, Germany  \n2 Leibniz University Hannover, Germany  \n{lena.john, soeren.auer, [oliver.karras}@tib.eu](oliver.karras}@tib.eu)[sushant.aggarwal@stud.uni-hannover.de](sushant.aggarwal@stud.uni-hannover.de)  \nAbstract Research Knowledge Graphs (RKGs) enable the structured representation of scientific knowledge, but their weakly enforced schemas make them prone to inconsistencies, particularly in how predicates are defined and used. Duplicate predicates, i.e., distinct identifiers expressing the same or highly similar relationships, introduce semantic redundancy, hinder reuse, and reduce RKG quality. While prior work has addressed duplicate detection for downstream tasks such as query answering or schema alignment, predicate redundancy as a data quality challenge, remains underexplored, particularly in terms of resolution, prevention, and semi-automated curator support.  \nIn this paper, we propose a framework for managing duplicate predicates in RKGs that covers detection, resolution, and prevention. The framework combines automated similarity-based methods with human validation and is designed for integration into the lifecycle of evolving, crowdsourced RKGs. We implement the framework in the context of the Open Research Knowledge Graph (ORKG) by extending its existing curation dashboard SciKGDash with embedding-based clustering, interactive inspection, and resolution actions such as merging and deleting.  \nWe evaluate the framework on the ORKG, where clustering reveals that up to 30% of predicates are potentially redundant. The analysis also shows recurring modeling patterns that lead to predicate redundancy, user-induced duplication, inconsistent identifier usage, and a lack of standardization in predicate naming and usage.  \nOur findings demonstrate that duplicate predicates arise from user behavior and interface design. Addressing this, requires combining automated methods with human-centered curation and preventive mechanisms. This work positions predicate redundancy as a central data quality challenge and provides a foundation for more systematic and proactive RKG curation.  \nKeywords: Duplicate Predicates · Knowledge Curation · Research Knowledge Graphs · Quality Assurance  \n2 L. John et al.  \n1 Introduction  \nKnowledge Graphs (KGs) model entities and their relationships in a semantically enriched graph structure, enabling both human and machine interpretability [11] . Research Knowledge Graphs (RKGs) have emerged as a key technology for representing scientific knowledge across various domains [31] .  \nTheir weakly enforced schemas and flexible nature support the integration of heterogeneous and evolving knowledge, but also introduce data quality challenges such as ambiguities, missing values, and duplicates [11,12,13] . Under the open world assumption (OWA) [10], (R)KGs are inherently incomplete [33] . Crowdsourced RKGs amplify these challenges by promoting decentralized and community-driven knowledge creation [22], which can lead to inconsistent modeling practices and evolving schemas [29] . Ensuring semantic consistency therefore requires continuous curation beyond initial data integration [12] .  \nA central yet often overlooked source of inconsistency lies in predicates. While entities can often be aligned to external references, predicates define how these entities are semantically related and thus shape the meaning of (R)KGs. Divergent modeling practices frequently result in duplicate predicates, i.e., distinct identifiers expressing the same or highly similar relationships [17,23] . These duplicates propagate semantic redundancy across all triples using them [8], ","cbCaihieYvV4IZGD","https://ap.wps.com/l/cbCaihieYvV4IZGD","pdf",3682611,1,18,"English","en",105,"# Introduction\n# Background: What is a duplicate predicate?\n# Related Work\n# Framework\n# Implementation\n# Preliminary Evaluation\n# Discussion of Findings\n# Conclusion and Future Work","[{\"question\":\"What problem does the paper address in research knowledge graphs?\",\"answer\":\"It addresses duplicate predicate redundancy in research knowledge graphs, where different identifiers express the same or highly similar relationships, causing semantic redundancy and hurting data quality.\"},{\"question\":\"How does the proposed framework identify duplicate predicates?\",\"answer\":\"The framework combines automated similarity-based methods with embedding-based clustering to surface potentially redundant predicates, followed by human validation.\"},{\"question\":\"How is the framework implemented and evaluated in the Open Research Knowledge Graph?\",\"answer\":\"The authors implement it in ORKG by extending SciKGDash with clustering, interactive inspection, and resolution actions such as merging and deleting. Evaluation shows clustering can reveal up to 30% of predicates potentially redundant and highlights recurring modeling patterns behind this issue.\"}]",1784197620,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"addressing-predicate-redundancy-in-research-knowledge-graphs","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/addressing-predicate-redundancy-in-research-knowledge-graphs/84678/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in research knowledge graphs?","Question",{"text":75,"@type":76},"It addresses duplicate predicate redundancy in research knowledge graphs, where different identifiers express the same or highly similar relationships, causing semantic redundancy and hurting data quality.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework identify duplicate predicates?",{"text":80,"@type":76},"The framework combines automated similarity-based methods with embedding-based clustering to surface potentially redundant predicates, followed by human validation.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the framework implemented and evaluated in the Open Research Knowledge Graph?",{"text":84,"@type":76},"The authors implement it in ORKG by extending SciKGDash with clustering, interactive inspection, and resolution actions such as merging and deleting. Evaluation shows clustering can reveal up to 30% of predicates potentially redundant and highlights recurring modeling patterns behind this issue.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]