Calculate String Similarity between two columns containing strings in two different dataframes

Viewed 144

I have one column in a particular dataframe (df1) containing certain number of observations. I want to calculate the highest similarity score between each observation of df1 and all the observations of another dataframe (df2) containing higher number of observations.

Example: df1-

Sr. No. Sentences
1. My name is Hitakshi
2. I am from US
3. "Hello! How are you?"
4. Stay right there

df2 -

Sr. No. Sentences
1. You are an idiot
2. Smart enough
3. My name is Aryan
4. What's up
5. Stay in
6. I am from US
7. Have patience
8. You are beautiful
9. My name is Hitakshi

Problem: I want to find the highest string similarity score between the first observation of df1 ('My name is Hitakshi') and all the observations of df2 (ignoring the punctuations). Likewise, for the second observation and so on.

Expected Output-

Sentences Highest Similarity Score
My name is Hitakshi 100
I am from US 100
"Hello! How are you?" 40
Stay right there 20

I know that I can use Jarowinkler distance, but how do I iterate each observation through all the observations of other column.

0 Answers
Related