How to use different error bars for grouped data of a pandas dataframe?

Viewed 159

I'm having some trouble with error bars in python. I'm plotting the columns on a pandas dataframe grouped, so on this example dataframe:

unfiltered = [0.975,0.964,0.689,0.974]
filtered = [0.954,0.932,0.570,0.960]
index_df = ["Accuracy", "Recall", "Precision", "Specificity"]
column_names = ["Unfiltered", "With overhang filter"]
df = pd.DataFrame(list(zip(unfiltered,filtered)),index=index_df,columns=column_names)

So my dataframe looks like this:

             Unfiltered  With overhang filter
Accuracy       0.975              0.954
Recall         0.964              0.932
Precision      0.689              0.570
Specificity    0.974              0.960

And I plot it with this following lines:

plt.style.use('ggplot')
ax = data_df.plot.bar(rot=0)
plt.show()

I get a figure like this:

linked image

Now I want to add error bars, but my problem is that I don't seem to be able to figure out how to get a different error value for each bar. I want to use the standard deviation and the values I have are different for each one of them (example: the std for both recalls shown is different). My problem is that if I add:

ax = data_df.plot.bar(rot=0, yerr=data_errors)

where data_errors is a list with the 8 standard deviations I get:

ValueError: The lengths of the data (4) and the error 8 do not match

It does work when data_errors has only 4 elements, but then it plots the same error bars for both accuracies, recalls, etc.

Can anyone help me to keep the data grouped by index like it is, but with different error bars for each value of the dataframe?

SOLUTION

Thanks to the user Quang Hoang I researched into sns.barplot. The solution to my problem was to create a dataframe (which I named data_df) like this:

      Indicator      Data       Class
0      Accuracy  0.966279  Unfiltered
1      Accuracy  0.981395  Unfiltered
2      Accuracy  0.989535  Unfiltered
3      Accuracy  0.975553  Unfiltered
4      Accuracy  0.961583  Unfiltered
5        Recall  0.954545  Unfiltered
...
35  Specificity  0.941176    Filtered
36  Specificity  0.953431    Filtered
37  Specificity  0.993865    Filtered
38  Specificity  0.946012    Filtered
39  Specificity  0.953374    Filtered

Followed by:

ax =  sns.barplot(x="Indicator", y= "Data",hue="Class", data=data_df, ci="sd")

This allowed me to create this figure: enter image description here where as you can see the error bars are different for each value, and also calculated automatically.

1 Answers

This might not be exactly what you're looking

data_df.stack().plot.bar(yerr=data_errors)
Related