I am a beginner in python trying to create a 2 component PCA plot, using pandas, sklearn.preprocessing, sklearn.decomposition, and Matplotlib.pyplot.
My data frame is very large, relating to the characteristics of different species of plant, with many variables (>100 columns), and I would like to compare the effect of one of the characteristics/columns (stem length) on the variance of the data. The column for stem length consists of floats, ranging in size from 0 to around 75cm.
I would like to plot a PCA comparing the variance of characteristics when stem length >40cm and stem length <40cm. However I have no idea how to proceed with this.
I have been using the following website as a guide for the PCA plot.
I have already written the following code:
import pandas as pd
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
df = pd.read_csv("plant_data.csv")
x = StandardScaler().fit_transform(x)
plt.style.use("seaborn-darkgrid")
pca = PCA(n_components=2)
principalComponents = pca.fit_transform(x)
principalDf = pd.DataFrame(data = principalComponents,
columns = ['principal component 1', 'principal component 2'])
finalDf = pd.concat([principalDf, df[['stem_length']]], axis = 1)
How do I set the conditions for the parameters to be stem_length >40 and stem_length <40?