scipy - Generate many two-sample t statistics using bootstrap

Viewed 309

I have the following pandas.DataFrame stored in df:

female income
0 23351.1
1 357.05
1 20385.2
0 12049.2

I want to use scipy.stats.bootstrap to calculate 1,000 t statistics comparing the mean income of men and women.

To generate a single t statistic, I would do:

# Import `stats`
from scipy import stats as st
# Get t-statistic
t = st.ttest_ind(df.loc[df['female'].eq(0), 'income'],
                 df.loc[df['female'].eq(1), 'income'])[0]

I was wondering if scipy.stats.bootstrap can be used to generate 1,000 such statistics.

I tried:

# Arrays for each group
f = df.loc[df['female'].eq(0), 'income'].values
m = df.loc[df['female'].eq(1), 'income'].values

# Wild attempt
st.bootstrap(data=[f, m], statistic=st.ttest_ind)

ValueError: `method = 'BCa' is only available for one-sample statistics

Here's a replicable example of df:

df = pd.DataFrame({'female': {0: 0, 1: 1, 2: 1, 3: 0, 4: 1, 5: 0, 6: 1, 7: 0, 8: 0, 9: 0},
 'income': {0: 23351.08,
  1: 357.04999,
  2: 20385.221,
  3: 12049.17,
  4: 20541.619,
  5: 13206.52,
  6: 17055.73,
  7: 60442.621,
  8: 4499.9902,
  9: 37663.039}})
0 Answers
Related