PySpark MLLib equivalent of sklearn DictVectorizer

Viewed 47

I am converting a sklearn build to a spark build. My data looks like this:

+-------+--------------------+
| source|             feature|
+-------+--------------------+
|7593648|Map(8495310 -> 9,...|
|7705409|Map(9517715 -> 6,...|
|7866217|Map(9159319 -> 9,...|
|8224541|Map(8224541 -> 7,...|
|8243189|Map(9159319 -> 7,...|
|8291454|Map(8727090 -> 6,...|
|8305210|Map(9266463 -> 57...|
|8420595|Map(9472594 -> 5,...|

feature column is a map of a string feature and an associated count. I was using a DictVectorizer in sklearn to generate feature vectors and then apply tfidf on them. I could not find an equivalent in MLLib. Any ideas?

0 Answers
Related