Spark MLLib Alternative to sklearn's TFIDF Vectorizer

Viewed 275

I am looking for an alternative to sklearn's TfIdfVectorizer in Spark's MLLib

As per Spark MLLib's documentation,

TF: Both HashingTF and CountVectorizer can be used to generate the term frequency vectors.

My question is, if I use MLLib's CountVectorizer for TF (since Spark's MLLib doesn't seem to have a TfIdfVectorizer), and use MLLib's IDF on top of that, would it be equivalent to sklearn's TfIdfVectorizer?

Or do I need to use something like a TfIdfTransformer over MLLib's CountVectorizer to achieve the same output as sklearn's TfIdfVectorizer? (which seems fishy to me)

NOTE : I am not talking about the data structure of the result but rather the equivalence of the algorithm implementations.

The limitation is that I cannot use HashingTF to keep the implementation equivalent to TfIdfVectorizer in sklearn

0 Answers
Related