set fixed value for stringIndexer encoded value Pyspark

Viewed 42

Given DataFrame that looks like this:

a      b       c.         d    
1      he     Null      fish
1      nb     canada     crab
2      ca     china     turtle
3      Null   chile      bird

Two task to accomplish:

  1. apply stringIndexer however want Null values to always be encoded to value 0.
  2. if column does not contain Null value I want it to have encoded value starting from 1.

My Initial approach was to use stringOrderType parameter in StringIndexer,

# added space to give lowest ascii code(32).
df = df.fillna(' unknown') 

cols = df.columns
indexers = {}
for col in cols:
    indexer = StringIndexer(inputCol=col, outputCol="ind_"+col, stringOrderType='alphabets')
    indexers[col] = indexer

When I apply this to my actual dataset some columns does not work as expected. Even if this works it cannot handle 2nd task.

What I've tried: Considering column 'c' for example:

labels = indexer.labels
# labels = [canada, chile,china, Null] this is order of index
# new_labels = [Null, canada, Chile, china]
indexer.transform(df, labels=new_labels)

In hoping to transform via new labels therefore Null gets mapped to 0 However obviously it does not have labels parameter hence error occurs. ValueError: Params must be a parammap but got list

0 Answers
Related