I'm working on a pipeline of manipulations on a pandas dataframe in class, and I'm wondering what the good steps are for concatenating some procedures one after the other - should I copy and recreate the original dataframe, or just change it in place?
According to the pandas documentation, working with views is not always recommended, and I'm not sure if this is the case here.
For example:
# first approach
import pandas as pd
class MyDataframe:
def __init__(self):
data = [['alice', 7], ['bob', 15], ['carol', 2]]
self.df = pd.DataFrame(data, columns = ['Name', 'Age'])
self.df_with_double_age_col = self._double_age()
self.df_with_double_age_col_and_first_letter = self._get_first_letter_of_name()
def _double_age(self):
self.df['double_age'] = self.df['Age'] * 2
return self.df
def _get_first_letter_of_name(self):
self.df['first_letter'] = self.df['Name'].str[0]
return self.df
In the preceding implementation, we obtain that self.df is equal to self.df_with_double_age_col and equal to self.df_with_double_age_col.
You can see that if you'll execute:
my_df = MyDataframe()
print(my_df.df.equals(my_df.df_with_double_age_col)) # True
I don't think it's a good situation (one thing with several names). I'll get a lot of aliasing if I add more processing steps. Alternatively, I was concerned about using only one dataframe and overwriting it at each step.
So, what do you see as any problem in that situation? I've included two additional optional implementations for that case below.
Overwrite (second approach):
import pandas as pd
class MyDataframe:
def __init__(self):
data = [['alice', 7], ['bob', 15], ['carol', 2]]
self.df = pd.DataFrame(data, columns = ['Name', 'Age'])
self.df = self._double_age()
self.df = self._get_first_letter_of_name()
def _double_age(self):
self.df['double_age'] = self.df['Age'] * 2
return self.df
def _get_first_letter_of_name(self):
self.df['first_letter'] = self.df['Name'].str[0]
return self.df
In-place change (without returns, third approach)):
import pandas as pd
class MyDataframe:
def __init__(self):
data = [['alice', 7], ['bob', 15], ['carol', 2]]
self.df = pd.DataFrame(data, columns = ['Name', 'Age'])
self._double_age()
self._get_first_letter_of_name()
def _double_age(self):
self.df['double_age'] = self.df['Age'] * 2
def _get_first_letter_of_name(self):
self.df['first_letter'] = self.df['Name'].str[0]
I believe that the last option is the most compact and elegant, but changing the given dataframe may be inconvenient and risky (in the context of SettingWithCopyWarning).