Fuzzy String Matching in Python using weightings

Viewed 113

I have Salesforce Records that I want to dedupe using fuzzy string matching techniques with weighting across different fields.

I want to set up scenarios such as weightings on specific columns in the row that increase or decrease the overall similarity metric. Essentially changing the weighting allow me to prioritize my columns at different levels.

I describe scenarios as a set of rules for how I want to compare records.

Below is an example data set:

First Last Name Email
Matt Metro name@example.com
Alex Two Three
Matthew Meos name@example.com

In this scenario we have 3 features for each row of data.

Each Feature has a weight of 10, this giving me a total score of 30

Feature Score
Fist Name 10
Last Name 10
Email 10
TOTAL 30

Thus an exact matching across all three fields would yield a 30 / 30 (i.e. 100% similarity score)

Now lets say I want to the weighting of email to be 3 times more powerful. The model should look like this:

Feature Score
Fist Name 10
Last Name 10
Email 30
TOTAL 50

Thus, an exact match with email, would hold significantly more weight in the similarity score.

I am trying to figure out the Python Packages that can help me achieve this dynamic weighting and the algorithm I should use for similarity.

For the Algorithm, I was thinking of using

  • Levenshtein Distance

For the Weightings of different fields between records:

  • Python Data Frame

What is the most performant way for achieving this this type of fuzzy string matching in Python?

0 Answers
Related