Setting min_impurity_decrease for a Regression Decision Tree

Viewed 217

I am thinking about the best way to set up a reasonable value for the min_impurity_decrease parameter for sklearn decision trees. It seems like one of the most important stopping criteria you can use, but the ideal parameter value strikes me as very ambiguous.

The issue seems a lot easier for classification trees, since gini impurity naturally ranges between 0 and 1. But for regression trees, the error metrics available to sklearn do not have any built in numeric range, so it seems like it's almost entirely determined by your data. The minimum amount of acceptable MSE reduction could vary wildly depending on your domain.

I know you can always grid-search these things but it'd be nice to have one less degree of freedom when searching for parameters.

What are the best decision criteria for setting this value?

0 Answers
Related