Different but the Same? An Event-Driven Approach to Determine Probabilities of Data Duplication

Authors: Heinrich, Bernd; Klier, Mathias; Obermeier, Andreas; Schiller, Alexander

Journal: MIS Quarterly (2025)

DOI: 10.25300/misq/2025/18178

<jats:p>The importance of data quality assessment, in general, and duplicate detection, in particular, has been recognized in both research and practice. Duplicates are known to cause critical problems in many domains, including customer relationship management, data management and data warehousing, fraud detection, production, and healthcare. Such duplicates are caused by events that are typically associated with data patterns in the records. For example, duplicates in customer databases caused by the event “relocation of a customer” characteristically exhibit dissimilar values for address-related features, while the values of the other features (e.g., name-related features) tend to be highly similar. Analyzing duplicate-related events and recognizing such patterns seems particularly promising for duplicate detection. However, existing approaches do not take advantage of this and neither consider events nor recognize their associated data patterns when detecting potential duplicates. In this paper, we introduce events as causes of duplicates and, on this basis, propose a novel probability-based approach for duplicate detection. Our approach assigns the probability of being a dupl…

View in Otero