Pyspark retrieve metrics (AUC ROC) from each submodel in CrossValidator

Viewed 537

How do I return the individual auc-roc score for each fold/submodel when using crossValidator.

The documentation indicates that collectSubModels=True should save all models rather than just the best or average, but after inspecting model.subModels I can't find how to print them.

The below example works just missing the model.subModels.aucScore

Desired Result would be each fold with their corresponding score like [fold1:0.85, fold2:0.07, fold3:0.55]

from pyspark.ml.feature import VectorAssembler
from pyspark.ml.classification import RandomForestClassifier
from pyspark.ml.tuning import CrossValidator, ParamGridBuilder
from pyspark.ml.evaluation import BinaryClassificationEvaluator

#Creating test dataframe
training = spark.createDataFrame([
    (1,0,1),
    (1,0,0),
    (0,1,1),
    (0,1,0)], ["label", "feature1", "feature2"])

#Vectorizing features for modelling

assembler = VectorAssembler(inputCols=['feature1','feature2'],outputCol="features")
prepped = assembler.transform(training).select('label','features')

#setting variables and configuring CrossValidator

rf = RandomForestClassifier(labelCol="label", featuresCol="features")
params = ParamGridBuilder().build()
evaluator = BinaryClassificationEvaluator()
folds = 3

cv = CrossValidator(estimator=rf,
estimatorParamMaps=params,
evaluator=evaluator,
numFolds=folds,
collectSubModels=True
)

#Fitting model
model = cv.fit(prepped)

#Print Metrics
print(model)
print()
print(model.avgMetrics)
print()
print(model.subModels)

>>>>>Return:
>>>>>CrossValidatorModel_3a5c95c6d8d2
>>>>>()
>>>>>[0.8333333333333333]
>>>>>()
>>>>>[[RandomForestClassificationModel (uid=RandomForestClassifier_95da3a68af93) with 20 trees], >>>>>[RandomForestClassificationModel (uid=RandomForestClassifier_95da3a68af93) with 20 trees], >>>>>[RandomForestClassificationModel (uid=RandomForestClassifier_95da3a68af93) with 20 trees]]
0 Answers
Related