I am trying to write a Python program which takes in multiple strings of text along with a phrase, and assigns each text string a score based on how much the concept discussed in the phrase actually appears in the text.
I want this to be a little more sophisticated than just a synonym finder, perhaps more similar to the kinds of searches Google, Google Scholar, or Semantic Scholar perform.
My current algorithm just tokenizes the phrase and runs individual synonym searches on each word, and pieces all the different combinations together to form new phrases. The results are pretty mediocre.
As an example, if I had the phrase "user-centered approach," I would like to be able to flag text strings that discussed things like human-computer interaction or human-focused design, even if the exact phrase is never used in the text.
Is there a way to achieve this or somethings similar in Python?