Store preprocessed data for later use via PyCharm

Viewed 110

After generating data for a machine learning algorithm using Keras, how do you save the LabelEncoder data?

The data is generated via the code:

from sklearn.preprocessing import LabelEncoder
from keras.utils import to_categorical

# Convert features and corresponding classification labels into numpy arrays
X = np.array(featuresdf.feature.tolist())
y = np.array(featuresdf.class_label.tolist())

# Encode the classification labels
le = LabelEncoder()
yy = to_categorical(le.fit_transform(y))

# split the dataset
from sklearn.model_selection import train_test_split

x_train, x_test, y_train, y_test = train_test_split(X, yy, test_size=0.2, random_state = 42)

From using Jupyter Notebook, one can save the data via:

### store the preprocessed data for use in the next notebook

%store x_train
%store x_test
%store y_train
%store y_test
%store yy
%store le

In PyCharm, I can successfully save the x_train data via:

savetxt(x_train.csv, x_train, delimiter=',')

However, this method does not work for the LabelEncoder.

I then tried using pickle via the code:

pickle.dump(le, open(filename, 'wb')

to save the encoded data. However, when I went to recall the data via:

LabelEnc = open(filename, 'rb')
le = pickle.load(LabelEnc)

I get the error: ...ValueError("Cannot load file containing pickled data " ValueError: Cannot load file containing pickled data when allow_pickle=False.

How can I correctly save and recall the le data?

This is a snippet of the data in PyCharm's "Show Variables" window:

enter image description here

0 Answers
Related