I'm confused why a DataFrame as a default function argument doesn't work quite as I expected. It puzzles me why I still have to specify it in the function call.
Basically I've created a function to merge dataframes to a main dataframe (which I set as the default argument), based on a key (like a LEFT JOIN)
Simplified Dataframes (Just the same student with different subject scores):
dfA = pd.DataFrame ({'student': ['A'],
'math_score': [50]})
dfB = pd.DataFrame({'student': ['A'],
'eng_score': [70]})
dfC = pd.DataFrame({'student': ['A'],
'sci_score': [80]})
Function (I want to get his subject scores one by one, by merging them to a main df)
def merge_selected(df_to_merge, main_df = dfA):
# LEFT JOIN to main_df
main_df = main_df.merge(right=df_to_merge, how='left', on=['student'])
return main_df
Using Function to Merging Twice:
dfA = merge_selected(dfB)
dfA = merge_selected(dfC)
Now here is what I don't get. I lost eng_score on the 2nd merge. Somehow when I assigned dfA to merge_selected a second time, it was removed.
student math_score sci_score
0 A 50 80
But if I specified main_df = dfA in the function call:
dfA = merge_selected(dfB, main_df = dfA)
dfA = merge_selected(dfC, main_df = dfA)
I don't lose eng_score and get all his scores:
student math_score eng_score sci_score
0 A 50 70 80
Essentially I've solved the issue, but was hoping someone could shed some light on this.
Why do I still have to specify main_df even though it is supposed to default to dfA?
Also, why do I lose eng_score if I do not specify main_df = dfA?