Column in DataFrame in Pandas with value 0

Viewed 125

I try to create 2 new columns in DataFrame in Pandas Python and the first column aa which shows average temperaturę is correct, nevertheless, the second column bb which should present temperature in City minus average temperature in all cities displays value 0??

Where is the problem? Did I correctly use lambda? Could you give me the solution? Thank you very much!

file["aa"] = file.groupby(['City'])["Temperature"].transform(np.mean)
display(file.sample(10))

file["bb"] = file.groupby(['City'])["Temperature"].transform(lambda x: x - np.mean(x))
display(file.head(10))
1 Answers

EDIT: Updated according to gereleth's comment. You can simplify it even more!

file['bb'] = file.Temperature - file.aa

Since we've already calculated the mean value in the aa column we can simply reuse this column to calculate the difference of the Temperature and aa column of each row by using pandas apply method like below:

file["aa"] = file.groupby(['City'])["Temperature"].transform(np.mean)
display(file.sample(10))
file["bb"] = file.apply(lambda row: row['Temperature'] - row['aa'], axis=1)
display(file.sample(10))

If you are looking to subtract the average of all cities temperature you can use mean on the column aa instead:

file["aa"] = file.groupby(['City'])["Temperature"].transform(np.mean)
display(file.sample(10))
avg_all_cities = file['aa'].mean()
file["bb"] = file.apply(lambda row: row['Temperature'] - avg_all_cities, axis=1)
display(file.sample(10))
Related