I am working on a dataset composed by 20060 rows and 10 columns and I am approaching decision tree regressor to make prediction.
My willing is to use RandomizedsearchCV in order to tune hyperparameters; my doubt is what to write in the dictionary as value for 'min_sample_leaf' and 'min_sample_split'.
My professor told me to rely on the database dimension but I don't understand how!
This is a code example:
def model(model,params):
r_2 = []
mae_ = []
rs= RandomizedSearchCV(model,params, cv=5, n_jobs=-1, n_iter=30)
start = time()
rs.fit(X_train,y_train)
#prediction on test data
y_pred =rs.predict(X_test)
#R2
r2= r2_score(y_test, y_pred).round(decimals=2)
print('R2 on test set: %.2f' %r2)
r_2.append(r2)
#MAE
mae = mean_absolute_error(y_test, y_pred).round(decimals=2)
print('Mean absolute Error: %.2f' %mae)
mae_.append(mae)
#print running time
print('RandomizedSearchCV took: %.2f' %(time() - start),'seconds')
return r_2, mae_
params= {
'min_samples_split':np.arange(), #define these two hypeparameter relying on database???
'min_samples_leaf':np.arange()
}
DT = model(DecisionTreeRegressor(), params)
Can anybody explain me?
Thank you very much