I have one column in a particular dataframe (df1) containing certain number of observations. I want to calculate the highest similarity score between each observation of df1 and all the observations of another dataframe (df2) containing higher number of observations.
Example: df1-
| Sr. No. | Sentences |
|---|---|
| 1. | My name is Hitakshi |
| 2. | I am from US |
| 3. | "Hello! How are you?" |
| 4. | Stay right there |
df2 -
| Sr. No. | Sentences |
|---|---|
| 1. | You are an idiot |
| 2. | Smart enough |
| 3. | My name is Aryan |
| 4. | What's up |
| 5. | Stay in |
| 6. | I am from US |
| 7. | Have patience |
| 8. | You are beautiful |
| 9. | My name is Hitakshi |
Problem: I want to find the highest string similarity score between the first observation of df1 ('My name is Hitakshi') and all the observations of df2 (ignoring the punctuations). Likewise, for the second observation and so on.
Expected Output-
| Sentences | Highest Similarity Score |
|---|---|
| My name is Hitakshi | 100 |
| I am from US | 100 |
| "Hello! How are you?" | 40 |
| Stay right there | 20 |
I know that I can use Jarowinkler distance, but how do I iterate each observation through all the observations of other column.