Definition
Soundex
Soundex is a phonetic matching algorithm that converts names into four-character codes based on how they sound. It helps identify names with different spellings but similar pronunciations, making it useful for name matching, record linkage, deduplication and data quality.
Soundex Calculator
Compare two names using the Soundex phonetic algorithm.
These values produce the same Soundex code.
Soundex keeps the first letter of a name and converts the remaining consonant sounds into a four-character phonetic code.
WinPure AdaptiveMatch™ can combine Soundex with other phonetic and fuzzy matching algorithms, thresholds and field-level rules to improve matching accuracy.
What Is the Soundex Algorithm?
Soundex is a phonetic matching algorithm designed to identify names that sound similar even when they are spelled differently. Instead of comparing names character by character, Soundex converts each name into a four-character phonetic code based on its pronunciation.
| Name | Soundex Code |
|---|---|
| Robert | R163 |
| Rupert | R163 |
| Smith | S530 |
| Smyth | S530 |
Because Robert and Rupert produce the same Soundex code, the algorithm identifies them as a potential phonetic match. The same happens with Smith and Smyth. This makes Soundex useful when working with datasets containing misspellings, alternative spellings or variations in people’s names. It has been widely used for searching records, identifying possible duplicates and linking information where exact text matching would miss a potential match.
Best Practice
Soundex is particularly suited to English-language names, but it has limitations. Two different names can sometimes generate the same code, while names that sound similar may occasionally produce different codes. For this reason, modern data matching systems often use Soundex alongside other phonetic and fuzzy matching techniques rather than relying on it alone.
How Does Soundex Work?
Soundex works by converting a name into a four-character phonetic code. Names that sound similar can produce the same code even when their spellings are different.
The algorithm follows four main steps:
1. Keep the first letter
The first letter of the name is retained. For example, with Robert, the letter R is kept.
2. Convert consonants into numbers
The remaining consonants are assigned numbers according to their sound groups:
| Number | Letters |
|---|---|
| 1 | B, F, P, V |
| 2 | C, G, J, K, Q, S, X, Z |
| 3 | D, T |
| 4 | L |
| 5 | M, N |
| 6 | R |
The letters A, E, I, O, U, H, W and Y do not receive a numeric value.
3. Remove duplicate adjacent sound codes
When adjacent letters produce the same numeric code, only one is retained. This prevents similar sounds from being counted more than once.
4. Create a four-character code
The first letter is combined with the first three numeric values. If fewer than three numbers remain, zeros are added until the code contains four characters.
Soundex Example: Robert
Although Robert and Rupert are spelled differently, they both produce R163. Soundex therefore considers them a phonetic match.
This is what makes Soundex useful for name matching and data deduplication: it can identify potential matches that an exact text comparison would miss.
Soundex Encoding Rules
Soundex uses a set of simple encoding rules to convert a name into a four-character phonetic code consisting of the first letter followed by three numbers.
1. Keep the First Letter
The first letter of the name is always retained.
For example: Robert → R
2. Convert Consonants into Numbers
The remaining consonants are grouped according to similar sounds:
| Soundex Code | Letters |
|---|---|
| 1 | B, F, P, V |
| 2 | C, G, J, K, Q, S, X, Z |
| 3 | D, T |
| 4 | L |
| 5 | M, N |
| 6 | R |
The letters A, E, I, O, U, H, W and Y are not assigned a numeric code.
3. Remove Repeated Sound Codes
If adjacent consonants belong to the same Soundex group, the repeated code is normally recorded only once. For example, Jackson contains consonants that can produce repeated sound codes. Soundex removes these duplicates rather than treating each one as a separate sound.
4. Handle Vowels, H, W and Y
Vowels and the letters H, W and Y do not receive Soundex numbers, but their treatment matters when determining whether consonant codes should be considered consecutive. In the commonly used American Soundex rules, vowels can separate identical consonant codes, while H and W do not necessarily break the sequence. This distinction is important because different simplified implementations of Soundex can occasionally produce different results.
5. Return Three Numbers
After applying the encoding rules, Soundex keeps the first three numeric codes. If there are more than three, the remaining codes are discarded.
6. Add Zeros When Required
If fewer than three numeric codes remain, zeros are added to the end.
For example: Lee → L000
The final Soundex key therefore always follows the same structure:
Best Practice
First letter + three digits
Examples include R163, S530 and L000.
These rules allow differently spelled names to be reduced to a common phonetic representation, making Soundex useful for identifying potential name variations and duplicate records that exact matching may overlook.
Soundex Examples
The easiest way to understand Soundex is to compare names with different spellings and see the phonetic codes they produce.
| Name A | Name B | Soundex A | Soundex B | Result |
|---|---|---|---|---|
| Robert | Rupert | R163 | R163 | Match |
| Smith | Smyth | S530 | S530 | Match |
| Ashcraft | Ashcroft | A261 | A261 | Match |
| Rubin | Ruben | R150 | R150 | Match |
| Jackson | Jaxen | J250 | J250 | Match |
| Robert | Roberts | R163 | R163 | Match |
| Smith | Schmidt | S530 | S530 | Match |
| Robert | Richard | R163 | R263 | No Match |
Robert vs Rupert
Robert and Rupert look quite different when compared character by character, but both produce:
R163
Soundex recognises that the important consonant sounds are similar, allowing the names to be identified as a potential phonetic match.
Smith vs Smyth
The difference between Smith and Smyth is the vowel used in the middle of the name. Because vowels are not assigned Soundex numbers, both names produce:
S530
This demonstrates why phonetic matching can find name variations that exact matching would miss.
Ashcraft vs Ashcroft
Both Ashcraft and Ashcroft produce: A261
Despite their spelling difference, Soundex reduces them to the same phonetic representation.
When the Codes Don’t Match
Soundex does not consider every name beginning with the same letter a match. For example:
Robert → R163
Richard → R263
Because the numeric portions differ, Soundex does not identify these names as a phonetic match.
What These Examples Show
Soundex can be particularly useful for finding alternative spellings, transcription differences and name variations within customer, CRM, historical and other datasets. However, sharing the same Soundex code does not mean two names definitely refer to the same person. For example, Smith and Schmidt both encode to S530 under American Soundex despite being distinct surnames. This is why Soundex is often most effective as one signal within a broader data matching process, alongside other attributes and matching algorithms.
How to Calculate a Soundex Code Step by Step
To calculate a Soundex code, start with the name and convert it into a four-character phonetic key.
Step 1: Keep the First Letter
Take the first letter of the name and keep it unchanged.
For example: Robert → R
Step 2: Convert the Remaining Consonants
Use the Soundex number groups:
| Number | Letters |
|---|---|
| 1 | B, F, P, V |
| 2 | C, G, J, K, Q, S, X, Z |
| 3 | D, T |
| 4 | L |
| 5 | M, N |
| 6 | R |
For Robert, the remaining relevant consonants are:
B → 1
R → 6
T → 3
So the code becomes: R163
Step 3: Remove Duplicate Sound Codes
If neighbouring consonants produce the same Soundex number, keep only one of them. This helps prevent repeated or closely related consonant sounds from being counted twice.
Step 4: Ignore Non-Coded Letters
The letters: A, E, I, O, U, H, W, Y – do not receive Soundex numbers.
They may still affect how duplicate consonant codes are treated, depending on the implementation.
Step 5: Keep the First Three Numbers
Soundex keeps a maximum of three numeric values after the first letter. If more than three codes are generated, any additional ones are ignored.
Step 6: Add Zeros if Necessary
If fewer than three numbers remain, add zeros until the Soundex code contains four characters.
For example: Lee → L000
Worked Example: Rupert
Now calculate the Soundex code for Rupert.
Keep the first letter: R
Ignore u
P → 1
Ignore e
R → 6
T → 3
The final code is: Rupert → R163
This is the same code as: Robert → R163
That is why Soundex identifies Robert and Rupert as a potential phonetic match. For larger datasets, this calculation is normally performed automatically. You can use the WinPure Soundex Calculator above to compare names instantly and see the phonetic codes produced.
Soundex for Name Matching
Soundex is particularly useful for name matching because names are often recorded differently across databases, CRM systems and historical records. Exact matching may treat two spelling variations as completely different values, even when they sound almost identical.
Soundex approaches the problem by comparing the phonetic representation of each name rather than relying only on its spelling.
| Name A | Name B | Soundex Codes | Exact Match | Soundex Match |
|---|---|---|---|---|
| Smith | Smyth | S530 / S530 | No | Yes |
| Robert | Rupert | R163 / R163 | No | Yes |
| Rubin | Ruben | R150 / R150 | No | Yes |
| Ashcraft | Ashcroft | A261 / A261 | No | Yes |
Share this article




