Handling NULL values in Spark StringIndexer

Viewed 6790

I have a dataset with some categorical string columns and I want to represent them in double type. I used StringIndexer for this convertion and It works but when I tried it in another dataset that has NULL values it gave java.lang.NullPointerException error and did not work.

For better understanding here is my code:

for(col <- cols){
    out_name = col ++ "_"
    var indexer = new StringIndexer().setInputCol(col).setOutputCol(out_name)
    var indexed = indexer.fit(df).transform(df)
    df = (indexed.withColumn(col, indexed(out_name))).drop(out_name)
}

So how can I solve this NULL data problem with StringIndexer?

Or is there any better solution for converting string typed categorical data with NULL values to double?

1 Answers
Related