I am trying to train the Radom Forest classifier.
Here is my code:
from sklearn.model_selection import GridSearchCV
from sklearn.ensemble import RandomForestClassifier
from sklearn.feature_extraction import DictVectorizer
import pickle
tr_features = pickle.load(open('./dataset/train_features.dump', 'rb'))
tr_labels = pickle.load(open('./dataset/train_labels.dump', 'rb'))
vectorizer = DictVectorizer()
vectorizer.fit(tr_features)
vectorized_features = vectorizer.transform(tr_features)
rf = RandomForestClassifier()
parameters = {
'n_estimators': [5, 50, 250],
'max_depth': [2, 4, 8, 16, 32, None]
}
cv = GridSearchCV(rf, parameters, cv=5)
cv.fit(vectorized_features, tr_labels)
After completion of training, I dumped the model with the joblib as follow:
import joblib
joblib.dump(cv.best_estimator_, './models/RF_model.pkl')
Later, when I load the model to run it on test data
import joblib
rf = joblib.load('./models/RF_model.pkl')
When I run the above line, the Jupuyter notebook's kernel restarts. When I checked the RF_model.pkl file size, it is 422.7MB.
I tried this solution, and passed the compress=3 argument to the joblib.dump() method.
joblib.dump(cv.best_estimator_, './models/RF_model.pkl',compress=3)
Even though the size changed to 43MB but still the kernel restarts and I cannot load the model.
Note: I am running the Jupyter notebook from a docker container.