Identity resolution is the process of determining that multiple data records, identifiers, or signals across different systems refer to the same real-world entity — whether a person, organisation, or device — and consolidating them into a unified, persistent representation. It combines probabilistic and deterministic matching algorithms, data enrichment, and graph-linking to resolve fragmented identities across first-party, second-party, and third-party data sources.

Content

  • Identity resolution as an enterprise practice emerged from direct marketing in the 1990s, when data brokers linked names and postal addresses across catalogue lists using surname-plus-address matching. The digital advertising ecosystem accelerated its sophistication dramatically: third-party cookies, device fingerprinting, and cross-site tracking created persistent identity graphs across billions of web sessions. Companies such as Acxiom, Experian, and LiveRamp built commercial identity resolution platforms connecting online and offline datasets at consumer scale. The discipline formalized around probabilistic record linkage theory (Fellegi-Sunter, 1969) adapted to real-time digital contexts.
  • Modern identity resolution operates in two modes. Deterministic resolution uses exact matches on strong identifiers — SHA-256 hashed email addresses, phone numbers, or social login tokens — to link records with near-certainty. Probabilistic resolution scores candidate pairs on weighted combinations of weaker signals: name tokens, device attributes, IP subnets, behavioural patterns, and temporal proximity. Machine learning models (gradient-boosted trees, Siamese neural networks) are trained on labelled identity pairs to optimise matching thresholds. The resulting identity graph is a property graph where nodes are identifiers and edges carry match confidence scores; connected components represent a single resolved entity. Real-time resolution pipelines process identity events at sub-10 ms latency to enable live personalisation.
  • Identity resolution matters across advertising, fraud prevention, healthcare, and security. Advertising platforms use it to measure cross-device campaign reach and frequency cap without double-counting. Banks and fintechs use it to detect synthetic identity fraud — where bad actors blend real and fabricated credentials — by catching statistical anomalies in how an identity’s components co-occur across time. Healthcare networks use it to link patient records across hospitals and insurers, preventing duplicate records that lead to medical errors. Cybersecurity teams use it to correlate attacker indicators across incidents, attributing campaigns to threat actors.
  • The 2024–2025 landscape is shaped by the deprecation of third-party cookies in Chrome (rolled out mid-2024), the growth of privacy-enhancing technologies, and tightening regulation under GDPR, CPRA, and India’s DPDP Act. Identity resolution vendors are pivoting to privacy-preserving approaches: clean rooms (where parties match hashed records without sharing raw data), universal IDs (such as UID2.0 and RampID based on consented email hashes), and on-device identity graph computation. Decentralised identity frameworks using W3C DIDs and Verifiable Credentials are positioned as a user-controlled alternative where individuals hold their own identity assertions rather than being passively resolved by data brokers.