I've a sample dataframe
pid = [1,2,3,4,5]; name = ['abc', 'def', 'bca', 'fed', 'pqr']; match_score = [np.nan, np.nan, np.nan, np.nan, np.nan]
sample_df = pd.DataFrame(zip(pid,name,match_score), columns=['pid', 'name', 'match_score'])
sample_df
| pid | name | match_score |
|---|---|---|
| 1 | abc | NaN |
| 2 | def | NaN |
| 3 | bca | NaN |
| 4 | fed | NaN |
| 5 | pqr | NaN |
And there's a name similarity score calculation method
from difflib import SequenceMatcher
SequenceMatcher(None, "abc", "bca").ratio()
>>> 0.666
How can I apply SequenceMatcher method to each row in the sample_df, so that I get
from difflib import SequenceMatcher
# comparing row1 with row2
print(SequenceMatcher(None, "abc", "def").ratio())
>>> 0.0
# comparing row1 with row3
print(SequenceMatcher(None, "abc", "bca").ratio())
>>> 0.66
# comparing row1 with row4
print(SequenceMatcher(None, "abc", "fed").ratio())
>>> 0.0
# comparing row1 with row5
print(SequenceMatcher(None, "abc", "pqr").ratio())
>>> 0.0
# Highest score for abc was 6.666
| pid | name | match_score |
|---|---|---|
| 1 | abc | 0.666 |
| 2 | def | NaN |
| 3 | bca | NaN |
| 4 | fed | NaN |
| 5 | pqr | NaN |
# comparing row2 with row1
print(SequenceMatcher(None, "def", "abc").ratio())
>>> 0.0
# comparing row2 with row3
print(SequenceMatcher(None, "def", "bca").ratio())
>>> 0.0
# comparing row2 with row4
print(SequenceMatcher(None, "def", "fed").ratio())
>>> 0.33
# comparing row2 with row5
print(SequenceMatcher(None, "def", "pqr").ratio())
>>> 0.0
# Highest score for def was 3.333
| pid | name | match_score |
|---|---|---|
| 1 | abc | 0.666 |
| 2 | def | 0.33 |
| 3 | bca | NaN |
| 4 | fed | NaN |
| 5 | pqr | NaN |
And so on:
| pid | name | match_score |
|---|---|---|
| 1 | abc | 0.666 |
| 2 | def | 0.333 |
| 3 | bca | 0.666 |
| 4 | fed | 0.333 |
| 5 | pqr | 0.000 |