Key Takeaways
- Identity resolution answers questions about people: which records, identifiers and events belong to the same individual.
- Entity resolution answers the same question for an entity: companies, suppliers, people data, then goes further by analysing how resolved entities relate to each other.
- The two use overlapping algorithms, but work differently. Identity graphs over-merge through shared identifiers; attribute matching trades precision against recall at the threshold.
- Identity intelligence is what an organisation can see once both are done well: who is who, and who is connected to whom.
- WinPure runs both inside the customer’s infrastructure, in Clean&Match Enterprise and through an API that embeds matching directly in your own applications.
Most data analysts and engineers have been doing identity and entity resolution for years under more familiar names: deduplication, record linkage, fuzzy matching, and merge and purge. While the two terms sound alike, and are often used interchangeably, their scope, objectives and outcomes differ.
Identity resolution deals with people data: employees, customers, patients in healthcare, citizens in government records and so on. Its aim is to establish which of the many variations of a person’s record belong to the same individual. Recognising that John Smith, J. Smith and Jonathan Smith, recorded with two phone numbers and three email addresses across different systems, are one person is identity resolution.
In contrast, entity resolution applies the same process beyond people. It resolves companies, suppliers, households, products, accounts and the other entities an organisation records, consolidates each into a single view, and identifies the relationships between them, such as the director who links two companies or the household that links several customers.
Knowing which of the two you need decides your approach to matching and consolidation: which fields count as evidence, how strict the match rules should be, and whether relationships between records need resolving. In this guide, we look at where disparate identities and entities cause problems in real data, touch on the matching logic behind each discipline, and show how WinPure Clean & Match Enterprise resolves both in one platform, with the WinPure API available for teams that want resolution built into their own systems.
Let’s get started.
Why it Matters to Understand Identity Resolution & Entity Resolution
In recent years, as organisations are realising the importance of cleaning and consolidating their CRM, ERP and operational systems data for business outcomes, identity resolution and entity resolution has become an important goal: For example, a CRM managers now know they need to consolidate multiple variations of their customer data into a trusted record that can then be used by their sales and marketing teams with more confidence. Data is no longer an entry in a spreadsheet – it’s an asset that demands the attention of both business and tech teams.
And most of that attention starts with customer data.
When Gartner’s Research Board for Global CIOs surveyed its members in April 2024, 8 of 14 respondents had a customer identity resolution solution in place, 11 of 12 named customer engagement and marketing as the main reason for implementing one, and half of those asked were running custom-built solutions. As projects like these extend beyond just getting a single view of the customer, identity and entity resolution begin to take on a new meaning – and knowing which part of the business they impact, matters.
1) They serve different business goals:
Because identity resolution centres on customer data, its most important uses sit in marketing: advanced personalisation, managing customers’ digital touchpoints and building customer intelligence for reporting. The Gartner figures reflect this, with customer engagement and marketing named as the main reason for implementing it.
Meanwhile, entity resolution supports the broader goal of organisational data quality, where deduplication, standardisation and governance apply to the overall data the business holds, from suppliers and accounts to products and policies.
2) They answer different questions:
The data behind customer identity resolution often sits in the CRM, where records can hold social profile links, digital activity and associations with other entities alongside basic contact details. Identity resolution uses that data to answer questions about one person:
- Who is John Smith?
- What is his job role, and where does he work?
- Which LinkedIn profile belongs to him?
- Is the John Smith now working at ABC the same person who used to work at XYZ?
Entity resolution answers broader questions about everything the organisation records, and about how those records connect:
- Are J & M Holdings Ltd and JM Holdings Limited the same company?
- Which suppliers share a bank account, registered address or director?
- How many distinct customers, accounts and households exist across CRM, ERP and billing?
- Which policyholder, claimant and broker records relate to the same party?
3. They are owned and funded by different teams
Marketing and customer experience teams usually own identity resolution, while data management, IT, compliance and operations teams own entity resolution. When the two are treated as one, organisations can end up buying overlapping capability from two budgets, or assuming a marketing tool has solved a data quality problem it was never scoped to address.
4. Entity resolution carries more complexity
Identity resolution deals with one entity type. Gartner’s definition of entity resolution and analysis covers individuals, products and other classes of data, along with the relationships among them, so a single dataset may need people, organisations, households and accounts resolved together. That wider scope needs more preparation, more careful matching configuration and stronger governance, and it should be resourced accordingly.
5. Identity resolution carries specific privacy obligations
Because identity resolution builds profiles of individuals, often for marketing, it sits squarely within data protection rules. Under the UK GDPR purpose limitation principle, organisations must be clear about their purposes from the start and can only reuse personal information for a new purpose if it is compatible with the original one. Linking customer records collected for service or billing into a marketing profile therefore needs a clear lawful basis, which is a different question from whether the records match.
While the above points help classify identity and entity resolution, it’s important to also know what kind of business goals or events trigger the need for either of the two. Knowing the business scope can help data analysts plan better – do they need a whole entity resolution activity or can a single customer view, achieved by identity resolution is more than enough. To help, here a few common scenarios, our data specialist teams at WinPure have witnessed that leads to either an identity or entity resolution project.
What Business Goals Create the Need for Identity and Entity Resolution – And How Should You Respond?
In a regular business meeting, you will rarely hear anyone talk about the data and say, ‘we need to resolve this entity,’ or ‘we need to consolidate customer information.’ Instead, you’re more likely to be asked to support the marketing team with a CRM clean-up, a personalisation programme or a business analytics project.
For example, in a B2B organisation holding thousands of supplier and vendor records, the request might be to improve billing accuracy or logistics. The conversation centres on the business outcome; while the data work is left for the data team to define.
And that definition starts once you open the pandora’s box (aka the data). What you find in the records tells you whether the objective needs identity resolution, entity resolution – or both. The following table depicts some of the most common business goals along with the type of resolution that works best.
Note: This mapping reflects what our teams have consolidated from customer projects over the years and can be used as a great starting point for scoping. The resolution and the tasks involved depend heavily on the data, the systems it sits in and what the organisation wants the result to support, so your project context may be entirely different from this.
| Business goal | Data challenge | Resolution needed | Associated data quality tasks |
|---|---|---|---|
| CRM clean-up or single customer view for personalisation | The same customer across web forms, imports and support systems; household members sharing emails and phone numbers | Identity resolution, with household grouping where the business communicates by address | Standardise emails and phone numbers, profile identifiers for shared or placeholder values, fuzzy and phonetic name matching, master record selection |
| Customer analytics and reporting | Duplicates inflating customer counts; B2B contacts not linked to their accounts | Identity and entity resolution | Deduplication, contact-to-account matching, golden records, persistent IDs for consistent reporting over time |
| Supplier billing, spend reporting or logistics | One supplier under trading and legal names across ERP, finance and procurement systems; missing registration numbers; outdated addresses | Entity resolution | Name and legal suffix standardisation, address parsing and validation, multi-attribute matching, merge and purge |
| CRM or ERP migration, or post-acquisition integration | Duplicates, conflicting formats and no shared key across several source systems at volume | Identity and entity resolution | Profiling, cleansing before load, cross-system matching, master record selection, persistent IDs |
| Patient or citizen record matching | The same person across services with missing or inconsistent identifiers; matching within a single facility can be as low as 80 percent | Identity resolution | Name, date of birth and address standardisation, deterministic and probabilistic matching, review of possible matches, audit trail |
| Compliance, KYC or fraud review | Connections through shared directors, addresses or bank accounts between records that are not duplicates | Entity resolution with relationship analysis | Resolving people and organisations together, relationship review, audit trail of match and merge decisions |
Most of these goals involves not just the ‘resolution’ component but also data quality – a task that only comes up when the data is inspected. For a migration or integration project, resolution and data quality become an engineering problem: how do you clean, standardise, compare records at scale so that the resulting matches can be trusted?
Comparing every record with every other grows quadratically, so systems group likely candidates before comparing them, as Binette and Steorts set out in their Sciences Advances review. Morever, missing and inconsistent values (messy data), need to be handled explicitly, which is a problem that impacts the quality of a record linkage, as confirmed by Fifield and Imai in their research on consolidating large scale records.
And this leads to a very interesting point of discussion – data quality.
Data Quality is the Foundation of a Well Executed Identity and Entity Resolution Process
In most projects, data quality is not the starting point of a discussion. The brief is simple: there’s an upcoming migration, or a business objective that requires the CRM to have reliable data. And everyone assumes the data is probably fine, until they start begin working on it. At that point, it becomes clear that the data needs more than a basic cleanup. Identity or entity resolution is simply not possible when other major issues like incomplete data, messy data, or years of legacy data has not been cleansed, transformed and deduped before it can be used.
Before records can be linked to the same person or organisation, they have to be complete, consistent and comparable – and that’s only possible when data quality is introduced as a task. The Office for National Statistics rightfully says, ‘a name recorded as Smith, Mr Bob will not match Robert Smith until the forenames are standardised and the name is split into separate forename and surname fields.’
This is one basic case of standardisation. The larger challenge comes from the quality challenges a real dataset contains:
- missing values in the fields you most need,
- misspelled names,
- addresses recorded in different formats or left out of date after a move,
- values entered in the wrong column, and
- placeholder entries such as dummy emails or strings of zeros.
Each of these has to be found through profiling, then corrected, standardised or flagged before any comparison between records can be relied on. But again, these are just basic challenges that happen with almost every data set. The deeper data quality issues are years and years of untouched legacy data or bloated data that needs more than just a basic cleanup.
In our 20+ years of working on some of the world’s most complex data sets, we’ve identified five major data quality challenges that need to be addressed before identity and entity resolution can take place. These are:
1. Data with no shared identifier across systems.
Each system issues its own keys. The CRM has a contact ID, the ERP has a customer number, billing has an account reference, and none of them was designed to link to the others – a problem that compounds when the project involves collaborating with third-party data, where the same data sources do not come with shared identifiers.
When the identifiers do not carry across, the connection between records has to be rebuilt from names, addresses, dates and other attributes, all of which carry their own errors.
The lack of shared identifiers between data sets causes a data consolidation challenge
See how you can match data, create unique IDs, and build master records that carry through.
2. Conflicting values for the same entity.
Two systems can agree that they hold the same customer or supplier and still disagree on the address, phone number, legal name or bank details. Resolving the duplicate is only half the job. Someone has to decide which source is trusted for which field, and without agreed rules the consolidated record can end up with the wrong address or the wrong payment details.
3. Structural inconsistency between sources.
One system stores a full name in a single field while another splits it into title, forename and surname. Addresses arrive as one line in one source and as separate components in another, trading and legal names share a field, dates follow different formats, and values end up in the wrong column entirely. These records cannot be compared field by field until their structure is aligned.
4. Data decay.
People move house, change names, change jobs and change email addresses. Companies rebrand, merge, move premises and dissolve. Historical records sit alongside current ones, and without dates or status flags there is often no way to tell which version reflects reality. Decay affects identity and entity resolution alike, because an outdated address can split one person into two records or join two different companies that once shared premises.
5. Duplicates that keep compounding.
A one-off clean-up improves the data on that specific day, but every web form, import, integration and manual entry after that point can create new duplicates and new inconsistencies. Unless new records are checked against the existing master as they arrive, or on a regular schedule, the quality gained from the clean-up gradually erodes and the business goal it supported weakens with it.
Key point to note: None of this is separate from an identity or entity resolution project. In fact, these data quality processes are part of the resolution process itself: for example, profiling shows you the severity of issues affecting your data, and the kind of cleaning methodology you will need to apply before you can match these groups together to resolve them. Once you have data you are confident with, is when you can apply the match logic to achieve the final outcome: identity or entity resolution.
Data Matching: The Engine Behind Identity and Entity Resolution
Whatever the business goal, every resolution project runs on the same underlying process: data matching. Matching compares records, scores how closely they agree and decides which ones belong together, whether that means one person, one household or one organisation. Resolution is the outcome, and matching is how it is achieved.
One of the most widely used forms in identity and entity resolution is fuzzy matching, which measures how similar two values are using string similarity measures such as Levenshtein distance or Jaro-Winkler, so that “Jonathon” and “Jonathan” can still be linked.
Fuzzy matching is only one type of match, though, and most projects draw on several others:
Exact or deterministic matching links records that agree precisely on a value. It works well for standardised identifiers such as email addresses, customer numbers or company registration numbers, and breaks down as soon as a value is mistyped or missing.
- Phonetic matching encodes names by how they sound, using algorithms such as Soundex or Double Metaphone, and catches variants such as “Smith” and “Smyth” that look different in writing.
- Token-based matching splits values into words before comparing them, so “Acme UK Holdings” and “UK Holdings Acme” are recognised as the same name despite the change in word order.
- Numeric and date matching compares values such as phone numbers, dates of birth or amounts, often with a tolerance for formatting differences or transposed digits.
- Probabilistic matching combines the evidence from several fields into a weighted score, so that agreement on a rare surname counts for more than agreement on a common first name.
- AI-based matching evaluates the whole record at once, which helps when values sit in the wrong field or when no single field is strong enough to decide a match on its own.
In practice, these types are combined, with each field assigned the comparison that suits the errors it typically contains:
| Field | Typical variation | Comparison that suits it |
|---|---|---|
| Personal names | Typos, initials, nicknames, spelling variants | Jaro-Winkler for spelling, phonetic encoding such as Double Metaphone, nickname libraries |
| Company names | Word order, legal suffixes, abbreviations | Token-based comparison after legal suffixes are normalised |
| Addresses | Abbreviated street types, missing flat numbers, reordered components | Component-level comparison after parsing, with local alignment methods such as Smith-Waterman-Gotoh for long strings |
| Emails and reference numbers | Case, spacing and system prefixes | Exact comparison after standardisation |
| Dates | Different formats, transposed day and month | Exact comparison with tolerance rules |
Now the typical approach to matching has been complex and frankly, a little inefficient.
Most data teams begin by building matching into the tools they already use: Python scripts on open-source record linkage libraries, matching rules written in SQL inside the data warehouse, or the deduplication features that come with their CRM. For many teams, this custom-built approach may seem like the right call, especially if they are working with just a few thousand records. But if they have years of legacy data, limited resources, and data that calls for more than simple matching logic – building it in-house may be tougher than it looks.
That said, before committing to a full build, it is worth measuring the whole cost.
Much of the work involves re-creating, testing, or extending the logic each time a new source or format arrives, keeping it fast as volumes grow, recording why every merge happened for audit and compliance, selecting master records and keeping identifiers stable across runs, and documenting it all so the knowledge does not sit with one or two people. That is a lot of engineering time! Spent on rudimentary mechanics, and less time on the business rules and domain judgement where a data team adds the most value.
A purpose-built matching platform can work alongside the team here, taking care of the algorithms, scale and audit trail so that data scientists and engineers can concentrate on defining what a match should mean for their business.
Deciding between build or buy?
Measure the costs, pros and cons between buying a matching software for resolution vs building one in-house.
Given that effort, it is understandable that some teams now reach for general-purpose AI tools as a shortcut, quickly pasting in sensitive customer records into a public chatbot to find duplicates without review. It feels faster, but it puts customer data outside the organisation’s control, hallucinates merges nobody can explain to an auditor, and ignores the business rules that decide what should be matched in the first place.
And the security consequences for using LLMs to match data is already widely published:
IBM’s 2025 Cost of a Data Breach Report found that high levels of shadow AI, meaning unapproved AI tools used by employees, added USD 670,000 to the global average breach cost, and that 97% of organisations reporting an AI-related security incident lacked proper AI access controls.
Customer and supplier records are exactly the data that matching projects handle, which makes these risks directly relevant to any team considering the shortcut. What the work actually needs is matching that combines the right algorithm for each field, runs inside the organisation’s own infrastructure and shows exactly why records were grouped.
When matching of that kind runs consistently across every source, with stable identifiers and the relationships between resolved records in view, the organisation gains a reliable answer to who is who and who is connected to whom. At WinPure we call that outcome identity intelligence.
Definition
What is Identity Intelligence in Data Management?
Identity intelligence is the capability to determine which records across all of an organisation’s sources refer to the same person or organisation, and how those resolved identities connect to one another. In this sense it is a data quality outcome built from records the organisation already holds, separate from identity security, which covers authentication and access control.
How WinPure Brings Data Quality, Matching and Resolution Together
With how data needs to be processed before we can resolve identities, one thing is clear: companies need a data quality solution that can handle matching and resolution as one process, deployed in the organisation’s own infrastructure, with results the team can explain.
WinPure Clean & Match Enterprise is built around that process.
Cleansing, standardisation, matching, master record creation, and identity and entity resolution run in one desktop-first platform with an embedded entity resolution engine, so data teams do not have to connect separate tools or send data outside their environment.
1. Profile and clean in the same project
The Data Quality Insights panel profiles every column for completeness, consistency, duplication and pattern anomalies. CleanAI™ then analyses the dataset and generates cleaning rules based on the field types and inconsistencies it detects, which analysts review, adjust and save for recurring runs. Word Manager and Pattern Manager handle organisation-specific standardisation, such as resolving “Ltd”, “Limited” and “Ltd.” to one form, while regex rules align phone numbers, postcodes and reference formats.

2. Match each field with the method that suits it
WinPure’s rule-based matching lets teams combine exact, fuzzy, phonetic and numeric comparison in one configuration, with a threshold and weight for each field. A typical identity configuration matches exactly on a standardised email or customer number, and falls back to fuzzy name comparison with exact date of birth and postcode where no identifier agrees. WinPure’s adaptive matching technology, AdaptiveMatch™, takes this further by reading what each column contains and recommending the approach that suits it, while experienced users keep full control over every setting. In WinPure’s internal testing, the latest matching engine improvements match one million records in around 15 seconds, which makes it much faster and easier to match records at scale.

3. Resolve entities and relationships with MatchAI™
MatchAI™ evaluates each record as a whole, so strong agreement across several fields can outweigh a poor score on one. It handles transposed names, addresses in the wrong column and companies recorded under an abbreviation in one system and a registered name in another, and it surfaces indirect relationships through shared phone numbers, addresses and partial name overlaps. Every grouping shows why the records were matched. In WinPure’s internal testing, MatchAI™ achieves 95 to 97% accuracy out of the box, without a configuration period.
4. Create master records and keep identities stable
SmartMaster AI™ scores each record in a match group on completeness, consistency and reliability and selects the golden record, using configurable, deterministic logic that data stewards can inspect and override. Persistent IDs give each resolved identity a stable identifier across processing runs, so new batches match against the established master without disturbing confirmed records.

5. Automate recurring runs and keep a full audit trail
Maintaining a resolved dataset involves maintaining a regular clean and deduplication schedule. WinPure has an automation module built to run cleaning, matching, and creating master record processes as scheduled, unattended batch jobs, so recurring imports and checked against the established master without having to rebuild the project over and over again. More importantly, every action done on the data is recorded in an audit log that captures time, user, and action performed, helping companies maintain their compliance processes without additional overheads.

What all this looks like on a live dataset
One of our customers, a direct-to-consumer retailer with around 850,000 customer records, needed both kinds of resolution in the same project. MatchAI™ resolved the records it could confirm as the same person and declined to merge records whose only shared value was an email address or phone number because, as our Director of Global Partnerships Kathryn Stevenson explained on the call, “the person could be different”. For this retailer, though, the household was the unit that mattered.
The configuration that worked kept every group MatchAI™ had resolved as its first rule, then added rule-based matching on names, email, account number, phone number and shipping region to capture household links, with the two levels kept separate for review. On the sample the team tested, the combined configuration identified 7,304 duplicates.
Watch this video to see how the customer resolved their data using WinPure.
One Engine, Two Ways to Work: the Desktop Platform and the WinPure API
Most resolution work in WinPure happens in Clean & Match Enterprise, where data analysts, data stewards and CRM administrators profile, clean, match and master records through the interface without writing code. Teams that also want matching inside their own applications, such as a CRM form that checks for an existing customer before creating a new one or a pipeline that deduplicates suppliers before they reach the warehouse, can use the WinPure API.
It embeds the same WinPure engine as a .NET library on Windows, Linux and macOS, processing data inside the customer’s own application, so new records can be checked against the trusted master that analysts maintain on the desktop.
Best Practice
If your team works through a data quality interface, Clean & Match Enterprise covers the whole process on its own. If you also build or maintain the applications that create records, the API lets you add matching at the point of entry. You can review the API documentation or talk to our team about integrating it.
A Real Case of API-Based Matching: HDL Companies
HdL Companies works with more than 700 US government agencies on revenue recovery and auditing. Its analysts compared leads from several sources in spreadsheets with VLOOKUPs, which offered no fuzzy logic and no consistent definition of a match. Because the team was already building its own lead management application, it integrated the WinPure API into that application: its own code handled cleaning, and the WinPure API matched leads and assigned each a match score based on the team’s custom fuzzy match logic.
Customer Quote
We generated over $1M just off that tool.
Jim Brashears, Software Development Manager, HdL Companies
HdL increased lead generation efficiency by more than 50%, generated over $1 million in new revenue through WinPure-based lead matching and improved its follow-up on tax delinquency.
Where to Start
Whether your goal needs identity resolution, entity resolution or both, the starting point is the same: understand what your data contains, decide what a match should mean for the business, and choose an approach that keeps data quality, matching and master records in one explainable process. Building in-house remains a valid route for teams with the time and appetite to maintain it. For teams that would rather spend that time on the business rules themselves, WinPure provides the cleansing, matching, entity resolution and master data capabilities in one platform, with the API available when resolution needs to run inside your own systems.
If you are scoping a resolution project and want to test your assumptions against real records, speak to our team about your entity resolution requirements. We will look at your sources, entity types and volumes with you and show where identity rules, MatchAI™ or the API fit.
Ready to resolve your identity and entity data?
Book Your Free 30 Days Trial.
It depends on what a wrong answer costs. A false merge in marketing is very different from one in healthcare or finance, and the best way to set thresholds is to test on a hand-checked sample.
The business defines what a match means, the data team builds the rules, stewards review borderline cases, and business owners check a sample before merging.
Disparate sources are not a barrier. WinPure imports each one into a single project, whether it is a CSV export from the CRM, an Excel file from billing, a JSON or XML extract, or a table from a SQL Server database, and works on a copy so the original systems stay untouched. You then choose the columns to match across the sources, such as name, email, phone number and postcode, and map them to each other even when each system names them differently. WinPure groups the records that belong to the same person and shows why they matched, so you can review the results before anything is merged.
It is recommended as a best practice to clean your data before matching it for entity resolution. In WinPure the cleaning happens in the same project as the matching, so there is no separate tool or export in between.
With rules where the business has them, and scoring where it does not. WinPure’s master record settings let teams prefer the record where a field such as VAT number is populated, or where a value is the longest or highest, or take the most populated record from a chosen table. Where no rule applies, SmartMaster AI™ scores each record in the group on completeness, consistency and reliability and selects the master.
While WinPure Clean&Match is a no-code platform, it is designed for data analysts and IT managers to program and structure match configurations around their specific requirements. The software offers full control so you can work on your data just the way you like it!
Text/CSV, Excel, JSON, JSONL, XML and SQL Server, and any custom data source. We also help with creating a connector if the need arises.
WinPure does not offer a SaaS version, and that is a deliberate choice. Clean & Match Enterprise is a locally deployed platform, so the organisation decides where it runs: on a desktop, on a server, in a virtual machine or air-gapped network, or in its own cloud environment on Azure or AWS. Wherever it is installed, data is processed inside the organisation’s own environment and is never sent to WinPure or a shared hosted service, which is what gives regulated and security-conscious teams control over their data. Teams that need embedded or automated processing inside their own systems can also use the WinPure API.
Annual licence, volume-based tiers, no per-record or API fees. It also covers which tiers include audit logging and automation.
First run within a day, personalised onboarding, and the 30-day free trial.
Share this article




