PySpark MLPC Multi Target classificaiton

Viewed 1206

PySpark 2.4.0

How to train a model which has multiple target columns?

Here is a sample dataset,

+---+----+-------+--------+--------+--------+
| id|days|product|target_1|target_2|target_3|
+---+----+-------+--------+--------+--------+
|  1|   6|     55|       1|       0|       1|
|  2|   3|     52|       0|       1|       0|
|  3|   4|     53|       1|       1|       1|
|  1|   5|     53|       1|       0|       0|
|  2|   2|     53|       1|       0|       0|
|  3|   1|     54|       0|       1|       0|
+---+----+-------+--------+--------+--------+

id, days and product are the feature columns. In order to train using PySpark ML - MLPC, i've converted the features into feature vectors.

Here is the code,

from pyspark.ml.linalg import Vectors
from pyspark.ml.feature import VectorAssembler

assembler = VectorAssembler(
    inputCols=['id', 'days', 'product'],
    outputCol="features")
output = assembler.transform(data)

and i've feature column as below,

+---+----+-------+--------+--------+--------+--------------+
| id|days|product|target_1|target_2|target_3|      features|
+---+----+-------+--------+--------+--------+--------------+
|  1|   6|     55|       1|       0|       1|[1.0,6.0,55.0]|
|  2|   3|     52|       0|       1|       0|[2.0,3.0,52.0]|
|  3|   4|     53|       1|       1|       1|[3.0,4.0,53.0]|
|  1|   5|     53|       1|       0|       0|[1.0,5.0,53.0]|
|  2|   2|     53|       1|       0|       0|[2.0,2.0,53.0]|
|  3|   1|     54|       0|       1|       0|[3.0,1.0,54.0]|
+---+----+-------+--------+--------+--------+--------------+

Now if i take each target columns as single label, i'll end up creating 3 models. But is there a way to convert all 3 targets(they are binary - 0 or 1) into labels.

For example if i take each target column separately then my MLPC layer will be like,

target_1 >> layers = [3, 5, 4, 2]
target_2 >> layers = [3, 5, 4, 2]
target_3 >> layers = [3, 5, 4, 2]

Since the target column contains only 0 or 1. Can i create a layer like below,

layers = [3, 5, 4, 3]

3 output for each target columns, they should give an output of 0 or 1 from every output neuron.

from pyspark.ml.classification import MultilayerPerceptronClassifier

trainer = MultilayerPerceptronClassifier(maxIter=100, layers=layers,blockSize=128, seed=1234)

I tried to combine all targets into single label,

assembler_label = VectorAssembler(
    inputCols=['target_1', 'target_2', 'target_3'],
    outputCol="label")
output_with_label = assembler_label.transform(output)

And the resulting data looks like,

+---+----+-------+--------+--------+--------+--------------+-------------+
| id|days|product|target_1|target_2|target_3|      features|        label|
+---+----+-------+--------+--------+--------+--------------+-------------+
|  1|   6|     55|       1|       0|       1|[1.0,6.0,55.0]|[1.0,0.0,1.0]|
|  2|   3|     52|       0|       1|       0|[2.0,3.0,52.0]|[0.0,1.0,0.0]|
|  3|   4|     53|       1|       1|       1|[3.0,4.0,53.0]|[1.0,1.0,1.0]|
|  1|   5|     53|       1|       0|       0|[1.0,5.0,53.0]|[1.0,0.0,0.0]|
|  2|   2|     53|       1|       0|       0|[2.0,2.0,53.0]|[1.0,0.0,0.0]|
|  3|   1|     54|       0|       1|       0|[3.0,1.0,54.0]|[0.0,1.0,0.0]|
+---+----+-------+--------+--------+--------+--------------+-------------+

When i tried to fit the data,

model = trainer.fit(output_with_label)

i got an error,

IllegalArgumentException: u'requirement failed: Column label must be of type numeric but was actually of type struct<type:tinyint,size:int,indices:array<int>,values:array<double>>.'

So, is there a way to handle data like this?

0 Answers
Related