NLP - obtain Beg Of Word > taxonomy matching rate with new text

Viewed 40

I'm struggling to understand the correct approach for the following problem:

  • I have some texts belong to same class (i.e. twitter text describing a particular product)
  • I want to generate a "taxonomy-beg-of-word"
  • I want to use the generated taxonomy to compare with next text in order to evaluate how much match the topic

I also know Bag of Words just creates a set of vectors containing the count of word occurrences in the document (i.e twitter texts), while the TF-IDF model contains information on the more important words and the less important ones as well.

But the point is, independently of BoX or TF-IDF, what is the correct approach to compare a new text against it? If I have for example this generated supervised product BoW form 3 different texts: enter image description here

I expect the dominant terms will be "shoes" and "modelY"

I don't get how to use this BoW to compare with new text and get a match rate (after stop-word removal and/or lemming stemming...)

For example:

  • "We sell new cars" > I'm expecting 0% match against BoW
  • "new shoes available" > I'm expecting some proportional rate match ("shoes" is an important match weithed to total token lengh)
  • "ModelY shoes on sale now!" > I'm expecting even higher rate as 2 important terms match the BoW

Any tips suggestion will be really appreciated. Many Thanks

0 Answers
Related