Situation
I’m trying to create a boxplot with individual and nested/grouped data. The dataset I use represents information for a number of households, where there is a distinction between 1-phase and 3-phase systems (#)
#NOTE Where the id appears only once, the household is single phased (1-phase) and duplicates are 3-phase system. Due to the duplicates, reading the csv-file via
pd.read_csv(..)will extend the duplicate's names (i.e.1,1.1and1.2).
Using the basic plot techniques delivers:
In [4]: VoltageProfileFile= pd.read_csv(dest + '/VoltageProfiles_' + str(PV_par['value_PV']) + '%PV.csv', dtype= 'float')
...: VoltageProfileFile.boxplot(figsize=(20,5), rot= 60)
...: plt.ylim(0.9, 1.1)
...: plt.show()
Out[4]:
The result is correct, but it would be clean to have only 1 tick representing 1, 1.1 and 1.2 or 5, 5.1, 5.2 etc.
Question
I would like to clean this up by using a ‘categorical’ boxplot, where values from duplicates (3-phase systems) are grouped under the same id. I’m aware that seaborn enables users to use the hue parameter: sns.boxplot(x='',hue='', y='', data='') to create categorical plots (Plotting with categorical data). However, I can’t figure out how to format my dataset in order to achieve this? I tried via pd.melt(..) function (cfr. pandas.melt), but the resulting format changes the order in which the values appear (*)
(*) Every id is accompanied by a length up to a reference point, thus the order of appearance on the x-axis must remain.
What would be a good approach to tackle this problem? Ideally, the boxplot would group 3-phase systems under one id and display different colours for 1ph vs. 3ph systems.
Kind regards,
Rémy


