Name Match V2
Introduction
Name Match is a difficult problem because there are a huge number of variations of the same name due to different reasons like OCR or human errors, alternate spellings, different order of names, skipping words, initials, salutations, etc.
To make this problem simple, this API from signzy not only gives a comprehensive match result and match score but also gives a match reason in response.
How to call the API
You will need to login before sending search request. You are required to pass the access token received from the login call, as authorization header in the Name match request.
Sample Curl
curl --location 'https://api.signzy.app/api/v3/matchers/nameMatchV2' \
--header 'accept: application/json' \
--header 'Content-Type: application/json' \
--header 'Authorization: <Token>' \
--data '{
"nameBlockv2": {
"name1": "<Name 1>",
"name2": "<Name 2>"
}
}'Input Parameters
PARAMETER | DESCRIPTION |
|---|---|
Content-Type | Contains the id parameter returned from the login step |
Authorization | application/json |
nameBlockv2 | Contains names to be matched |
name1 | First name to be matched |
name2 | Second name to be matched |
Sample Response
{
"result": {
"name1_vs_name2_matchScore": 1,
"name1_vs_name2_matchResult": "Direct match",
"name1_vs_name2_matchReason": "The given names are identical"
}
}Response Parameters
PARAMETER NAME | DESCRIPTION |
|---|---|
name1_vs_name2_matchScore | Name Match score between 0 to 1 |
name1_vs_name2_matchResult |
|
name1_vs_name2_matchReason | Message output for explaining the reason for the score as a means of text based feedback on what happens inside the code |
Details on Output parameters
matchResult
Gives a single class denoting how well the names match. These classes are a grouped representation of a range of scores as categorized below.
Table 1: Score Categorization
MATCHRESULT (V1) | MATCHRESULT (V2) | MATCHSCORE RANGE |
|---|---|---|
Direct Match | Direct Match | 1 |
Partial Match | Good Partial Match | 0.85 - 0.99 |
| Moderate Partial Match | 0.60 - 0.84 |
| Poor Partial Match | 0.34 - 0.59 |
| No Match | 0 - 0.33 |
No Match | No Match | 0 |
name1_vs_name2_matchScore
Gives the raw matching score as a number between 0 to 1. A score of 0 means totally different names and a score of 1 means exactly the same names.
name1_vs_name2_matchReason
Gives a transparent reason why the Name Match model thought the input names are similar or different. This is directly inferred from the modular structure of our algorithm.
*This feature is still under development and we are working on improving it further.
Name 1 | Name 2 | matchResult | matchScore |
|---|---|---|---|
RAVINDRA PRATAP SINGH | RAVINDRA PRATAP SINGH | Direct match | 1.00 |
These names are exactly same, so we get a perfect score | | | |
RAVINDRA PRATAP SINGH | RAVINDRA PRATAP | Good partial match | 0.85 |
Here, one name (SINGH) is missing but the rest of the names match perfectly. So we are pretty sure it"s the same person | | | |
RAVINDRA PRATAP SINGH | RAVINDRA PRATAP S | Good partial match | 0.85 |
Same as above, here we also have an initial “S” for the missing word, so it"s a good partial match | | | |
RAVINDRA PRATAP SINGH | R P SINGH | Moderate partial match | 0.70 |
Here, 2 long names are missing (RAVINDRA and PRATAP) so the score is lower but since the initials match, we are moderately sure it"s the same person | | | |
RAVINDRA PRATAP SINGH | RAVINDRA | Moderate partial match | 0.70 |
Same as above, 2 names are missing | | | |
RAVINDRA PRATAP SINGH | SINGH RAVINDRA PRATAP | Good partial match | 1 |
Here all three names are present in both but in a different order. But overall it"s a good match, we are not penalizing this currently | | | |
RAVINDRA PRATAP SINGH | RAVINDRA PRATAP SINGH S O DH | Good partial match | 0.90 |
Here, there are a few extra words (S O DH) and Name Match recognizes it stands for caregiver name (as in Son of DH) and so doesn"t penalize it highly. Overall result is a good match | | | |
RAVINDRA PRATAP SINGH | RAVINDRA T SINGH | Poor partial match | 0.55 |
Here, although 2 names are similar but the third initial “T” is completely different which means there is a chance it is a different person so Name Match penalizes it higher | | | |
RAVINDRA PRATAP SINGH | RAVINDRA THAKUR | Poor partial match | 0.38 |
Here, “Thakur” is completely different than “Pratap” and “Singh” which means it is highly likely this is a different person, so the final score is very low | | | |
RAVINDRA PRATAP SINGH | NARENDRA SINGH | Moderate partial match | 0.61 |
Here, “Singh” completely matches in both names and “NDRA” from both names matches which increases the score but overall it is still a partial match | | | |
SHINDE RAHUL MARUTI | S R MEDICO | Poor partial match | 0.41 |
Here, the names are completely different. However, S and R match with the initials of SHINDE and RAHUL so there is some score given. Overall it"s still a poor match | | | |
LALAN PRASAD | MAHTO RAMBHAJAN LALANPRASAD | No match | 0.20 |
Here, as humans we can probably see that there is some chance that these two names are of the same person because of the part “LALANPRASAD” in the names but Name Match penalizes it highly due to missing space and words | | | |
ATULKUMAR MELABHAI PARMAR | ATULBHAIMELABHAIPARM | No match | 0.14 |
Although we can see there are parts of this name which tells us this might belong to the same person, Name Match penalizes it highly. | | | |
Version History
V1
Similarity between two names was calculated by using distance algorithms which did not care for subtle nuances in Names and treated all variations in alphabets, characters, names equally.
V2 Changes
nameMatchV2 runs the new improved version which enacts the improvements described below
Better score categorization
V1 only had "Direct Match", "Partial Match" and "No Match".
V2 breaks it up into:
- Direct Match
- Good Partial Match
- Moderate Partial Match
- Poor Partial Match
- No Match
This enables more granular score categorization. For example, let's see 2 different cases and the difference between V1 and V2 response
Input: 'Tonmoy Sahu' and 'Aditya Sahu'
| V1 response | V2 response |
|---|---|---|
name1_vs_name2_matchScore | 0.50 | 0.60 |
name1_vs_name2_matchResult | Partial Match | Moderate partial match |
Input: 'Tonmo Sahu' and 'Tonmoy Sahu'
| V1 response | V2 response |
|---|---|---|
name1_vs_name2_matchScore | 0.92 | 0.80 |
name1_vs_name2_matchResult | Partial Match | Good partial match |
Modules to captures nuances of Names
We have introduced modules that take care of groups of user scenarios as shown below. These modules also enable us to return specific matchReason based on the module it hits and give appropriate penalties based on the scenario.
Scenario | Module |
|---|---|
Exact names from two sources | Exact match |
One of the names is reversed and many times has a comma | Sequence |
Users name in One ID has a middle name the other doesn't | Word missing |
Users name in One ID has initial of one name the other doesn't (Tonmoy J Borah - Tonmoy Borah) | Initials |
The OCR of one ID card has removed spaces b/w a name hence the name from one ID card has no spaces | Space |
The user's name in one ID card has a different spelling but the same pronunciation as some names in Hindi and English are spelled differently. Like RAJEEV and RAJIV | Spelling substitution |
Names are taken from an ID with Salutations Like Mr and Dr | Pretreatment |
Two names are Naval Kishore and NKishore or some letters are missing | Subset |
Similar spellings of names (agarwala - agarwal, mahammad- mohd) in certain cultures | Common Word Substitution |
Sometimes names have caregiver names (eg. Aditya Jain S/O Ramesh and Aditya Jain) | Care of module |
We plan to continuously improve and add support for new customer expectations and real world scenarios which will be bundled into different modules and either be included into the Default version or be kept separate.
Status Codes
CODE | MESSAGE |
|---|---|
400 | Error: Name cannot be blank or spaces or anything except string |
500 | Internal Server Error |
Getting help
Please feel free to contact us if you have any questions, require clarification, or have ideas for how to make the documents or any of our services better.
You can reach out to us at [email protected].