I am looking for an alternative to sklearn's TfIdfVectorizer in Spark's MLLib
As per Spark MLLib's documentation,
TF: Both HashingTF and CountVectorizer can be used to generate the term frequency vectors.
My question is, if I use MLLib's CountVectorizer for TF (since Spark's MLLib doesn't seem to have a TfIdfVectorizer), and use MLLib's IDF on top of that, would it be equivalent to sklearn's TfIdfVectorizer?
Or do I need to use something like a TfIdfTransformer over MLLib's CountVectorizer to achieve the same output as sklearn's TfIdfVectorizer? (which seems fishy to me)
NOTE : I am not talking about the data structure of the result but rather the equivalence of the algorithm implementations.
The limitation is that I cannot use HashingTF to keep the implementation equivalent to TfIdfVectorizer in sklearn