I am conducting PCA on a spam email dataset, everything is fine up until the point where I want to graph the principal components against each other, pc1vspc2, pc1vspc3 and pc2vspc3. The scatter plots are running okay, but I want to display the spam data points on top of the nonspam data.
I have searched up and down for a way to do this, but can't seem to find any method that works!
#Seperating Feautures
X = df.iloc[:,:54]
#Seperating Target, changing 0's to non-spam & 1's to spam
Y = df['Spam_Indicator'].values.tolist()
for i in range(len(Y)):
if Y[i] == 1:
Y[i] = 'Spam'
else:
Y[i] = 'Non-spam'
Y = np.asarray(Y)
#no of principal components
N = 3
col_numbering = [str(x) for x in range(1,N + 1)]
#Applies PCA reducing from 54 to N dimensions
pca = PCA(n_components = N)
X_red = pca.fit_transform(X)
X_red = pd.DataFrame(data = X_red, columns = col_numbering)
#Prints the components, explained variance and explained variance ratio
#print('Components:',pca.components_)
print('Explained Variance:' ,pca.explained_variance_)
print('Explained Variance Ratio:' ,pca.explained_variance_ratio_)
plt.figure(figsize=(20,10))
plt.subplot(1,3,1)
sns.scatterplot(x = '1', y = '2', data = X_red, hue = Y,
alpha = .75, hue_norm = (0.7))
plt.subplot(1,3,2)
sns.scatterplot(x = '1', y = '3', data = X_red, hue = Y,
alpha = .75, hue_norm = (0.7))
plt.subplot(1,3,3)
sns.scatterplot(x = '2', y = '3', data = X_red, hue = Y,
alpha = .75, hue_norm = (0.7))
plt.show()
Here is an image of what I have so that you know better what it is I'm asking. Seaborn Scatter Plot