I have Salesforce Records that I want to dedupe using fuzzy string matching techniques with weighting across different fields.
I want to set up scenarios such as weightings on specific columns in the row that increase or decrease the overall similarity metric. Essentially changing the weighting allow me to prioritize my columns at different levels.
I describe scenarios as a set of rules for how I want to compare records.
Below is an example data set:
| First | Last Name | |
|---|---|---|
| Matt | Metro | name@example.com |
| Alex | Two | Three |
| Matthew | Meos | name@example.com |
In this scenario we have 3 features for each row of data.
Each Feature has a weight of 10, this giving me a total score of 30
| Feature | Score |
|---|---|
| Fist Name | 10 |
| Last Name | 10 |
| 10 | |
| TOTAL | 30 |
Thus an exact matching across all three fields would yield a 30 / 30 (i.e. 100% similarity score)
Now lets say I want to the weighting of email to be 3 times more powerful. The model should look like this:
| Feature | Score |
|---|---|
| Fist Name | 10 |
| Last Name | 10 |
| 30 | |
| TOTAL | 50 |
Thus, an exact match with email, would hold significantly more weight in the similarity score.
I am trying to figure out the Python Packages that can help me achieve this dynamic weighting and the algorithm I should use for similarity.
For the Algorithm, I was thinking of using
- Levenshtein Distance
For the Weightings of different fields between records:
- Python Data Frame
What is the most performant way for achieving this this type of fuzzy string matching in Python?