I have the following pandas.DataFrame stored in df:
| female | income |
|---|---|
| 0 | 23351.1 |
| 1 | 357.05 |
| 1 | 20385.2 |
| 0 | 12049.2 |
I want to use scipy.stats.bootstrap to calculate 1,000 t statistics comparing the mean income of men and women.
To generate a single t statistic, I would do:
# Import `stats`
from scipy import stats as st
# Get t-statistic
t = st.ttest_ind(df.loc[df['female'].eq(0), 'income'],
df.loc[df['female'].eq(1), 'income'])[0]
I was wondering if scipy.stats.bootstrap can be used to generate 1,000 such statistics.
I tried:
# Arrays for each group
f = df.loc[df['female'].eq(0), 'income'].values
m = df.loc[df['female'].eq(1), 'income'].values
# Wild attempt
st.bootstrap(data=[f, m], statistic=st.ttest_ind)
ValueError: `method = 'BCa' is only available for one-sample statistics
Here's a replicable example of df:
df = pd.DataFrame({'female': {0: 0, 1: 1, 2: 1, 3: 0, 4: 1, 5: 0, 6: 1, 7: 0, 8: 0, 9: 0},
'income': {0: 23351.08,
1: 357.04999,
2: 20385.221,
3: 12049.17,
4: 20541.619,
5: 13206.52,
6: 17055.73,
7: 60442.621,
8: 4499.9902,
9: 37663.039}})