I was searching for machine learning examples to look at and understand and I stumbled upon this example: https://www.kaggle.com/saulalquicira/model-evaluation-using-cross-val-score-and-kfold
I understand everything in the code except for this part:
labelencoder_X = LabelEncoder()
X[:,2] = labelencoder_X.fit_transform(X[:,2])
ct = ColumnTransformer([("cp", OneHotEncoder(), [2])], remainder = 'passthrough')
X = ct.fit_transform(X)
ct = ColumnTransformer([("restecg", OneHotEncoder(), [9])], remainder = 'passthrough')
X = ct.fit_transform(X)
ct = ColumnTransformer([("slope", OneHotEncoder(), [15])], remainder = 'passthrough')
X = ct.fit_transform(X)
ct = ColumnTransformer([("ca", OneHotEncoder(), [18])], remainder = 'passthrough')
X = ct.fit_transform(X)
ct = ColumnTransformer([("thal", OneHotEncoder(), [22])], remainder = 'passthrough')
X = ct.fit_transform(X)
I understand what every individual keyword does, but why are we using this on values that are already numerical in nature, I thought we do this on categorical Data that is alphabetical in nature in order to transform it to numerical binary values that machine learning algorithms can understand. here is how the Dataset looks:
