I have data like this:
| Term | Value|
| -------- | -----|
| Apple | 100 |
| Appel | 50 |
| Banana | 200 |
| Banan | 25 |
| Orange | 140 |
| Pear | 75 |
| Lapel | 10 |
Currently, I am using the following code:
matches = []
for term in terms:
tlist = difflib.get_close_matches(term, terms, cutoff = .80, n=5)
matches.append(tlist)
df["terms"] = matches
The output is like this
| Term | Value|
| --------------------- | -----|
| [Apple, Appel] | 100 |
| [Appel, Apple, Lapel] | 50 |
| [Banana, Banan] | 200 |
| [Banan, Banana] | 25 |
| [Orange] | 140 |
| [Pear] | 75 |
| [Lapel, Appel] | 10 |
This code isn't really helpful. My desired output is something like:
| Term | Value|
| -------- | -----|
| Apple | 150 |
| Banana | 225 |
| Orange | 140 |
| Pear | 75 |
| Lapel | 10 |
The main issue is that the lists aren't in the same order, and often there is only one or two words of overlap in the lists. For example, I might have
- [apple, appel]
- [appel, apple, lapel]
Ideally, I would like to have both these return "apple", because that has the highest value of the overlapping terms.
Is there a way to do this?