When calculating a mean in two different ways (on a dataframe and on the same pivoted dataframe) I expect the outcomes to be identical. However, they appear to differ. Am I missing something?
Here's the dataset:
import pandas as pd # pandas version is 1.3.4
df = pd.read_csv(
'https://data.rivm.nl/covid-19/COVID-19_aantallen_gemeente_per_dag.csv',
usecols = ['Date_of_publication', 'Municipality_code', 'Municipality_name', 'Province', 'Total_reported', 'Hospital_admission', 'Deceased'],
parse_dates = ['Date_of_publication'],
index_col = ['Date_of_publication'],
sep = ';'
).dropna()
df.tail()
I would like to calculate a mean per Date_of_publication of the column Total_reported.
Method 1:
df.Total_reported.groupby(df.index).mean()
Method 2:
df_pivot = pd.pivot_table(
df.reset_index(),
values='Total_reported',
index='Date_of_publication',
columns='Municipality_name'
)
df_pivot.mean(axis=1)


