Should sklearn.ensemble.AdaBoostClassifier be relying on duplicate estimators (weak learners)?

Viewed 213

While analyzing the errors (misclassifcations) of an sklearn.ensemble.AdaBoostClassifier using DecisionTreeClassifier stubs as the base estimators, I have found that there a large number of duplicated estimators in the ensemble. Is it typical to have this amount of redundancy? Details below.

Version information

python: 3.6.5 sklearn version: 0.20.2

Here is the setup with the number of estimators and learning rate identified using GridSearchCV.

    number_estimators = 501
    bdt= AdaBoostClassifier(DecisionTreeClassifier(max_depth=1), 
    algorithm=algorithm_choice, n_estimators=number_estimators,
                   learning_rate = 1)

For reference a partial list of important features:

    important_features = utils.display_important_features( bdt.feature_importances_, X_train)
    Coefficient Values
    delta_sos   0.14770459081836326
    delta_win_pct   0.14171656686626746
    delta_dol   0.06786427145708583
    delta_fg_pct   0.0658682634730539
    delta_col   0.06387225548902195
    ...

Examining the estimators:

    feature_dict={}
    threshold_dict = {}
    for stub_estimator in bdt.estimators_:
        stub_tree = stub_estimator.tree_
        stub_feature_index = stub_tree.feature[0]
        stub_feature = X_train.columns[stub_feature_index]
        if stub_feature in feature_dict:
           feature_dict[stub_feature] +=1
           threshold_dict[stub_feature].append(stub_tree.threshold[0])
        else:
           feature_dict[stub_feature] = 1
           threshold_dict[stub_feature] = []
           threshold_dict[stub_feature].append(stub_tree.threshold[0])

     feature_dict
     {'delta_srs': 31,
      'delta_win_pct': 71,
      'delta_sos': 74,
      'delta_sag': 13,
      'delta_dol': 34,
       ... }

Let's consider the 'delta_win_pct' feature. Of the 501 estimators in the ensemble, the boosting algorithm has selected 71 of the estimators to be based upon the 'delta_win_pct'. As a note, the percentage of 'delta_win_pct' estimators, 71/501 = 14.17%, is the value associated with the important feature attribute associated with the classifier as noted above.

Of the 71 'delta_win_pct' estimators the thresholds were also collected in the code block above and can be accessed via:

   delta_win_partitions = threshold_dict['delta_win_pct']
   delta_win_partitions = [i*100 for i in delta_win_partitions]
   df_win = pd.DataFrame(columns=['Delta_Win_Pct','Y'])
   df_win['Delta_Win_Pct'] = delta_win_partitions
   df_win['Y'] = 1
   df_win.sort_values(by='Delta_Win_Pct', inplace=True)

   print("Number of unique partitions= 
        ",df_win['Delta_Win_Pct'].unique().shape[0], " out of ", 
         df_win.shape[0], ' estimators')

   splot = sns.scatterplot(x='Delta_Win_Pct', y='Y', data=df_win)
   splot.figure.set_size_inches(20,6)
   splot.set_title('Partitions of Delta Win Percentage')  

   ------------
   Number of unique partitions=  27  out of  71  estimators

enter image description here

The plot depicts the unique 27 thresholding values used to partition the 'delta_win_pct' feature in the ensemble of estimators ranging from -25.1% to 22.2%.

The issue is that 71-27= 44, or 61.98% of the 'delta_win_pct' tree stubs use duplicate threshold values contained in the plot.

Isn't this inefficient? Shouldn't the estimator weights reflect the contribution of the estimators for the classifier and not the rely on duplicate estimators to garner more votes?

Should the algorithm ensure that no duplicate threshold values are selected for the same feature based estimators?

With 71 estimators being used for the 'delta_win_pct' feature and eliminating duplicate threshold values, partitions could occur at many more intervals yielding finer granularity in carving up the feature. Or, if the finer granularity does not help in the construction of the decision tree stub then eliminating the duplicate values would reduce the ensemble size.

Is this thinking correct?

Have other practitioners attempted to minimize the duplication of estimators in the ensemble?


Update 1/29/2019

The number of duplicated estimators is certainly driven by the number of requested estimators and how a decision tree feature stump can be split via Entropy or Gini. The splitting of the feature stump is performed independently of other stumps in the forest in that no list of previous thresholds for a specific feature stump is maintained. The thresholds appear to be determined solely by the splitting algorithm.

What I have found empirically is that a high number of duplicate tree stubs is indicative of requesting too many estimators. By simply reducing the requested estimators, duplicates are reduced without loss in the predictive power of the algorithm.

0 Answers
Related