Display additional values in holoviews sankey labels or hover information

Viewed 577

I would like to find a way to modify the labels on holoviews sankey diagrams that they show, in addition to the numerical values, also the percentage values.

For example:

import holoviews as hv
import pandas as pd
hv.extension('bokeh')


data = {'A':['XX','XY','YY','XY','XX','XX'],
        'B':['RR','KK','KK','RR','RK','KK'],
        'values':[10,5,8,15,19,1]}

df = pd.DataFrame(data, columns=['A','B','values'])
    
sankey = hv.Sankey(df)

For 'From' label 'YY' which is 'YY - 8' change this to 'YY - 8 (13.7%)' - add the additional percentage in there.

I have found ways to change from the absolute value to percentage by using something along the lines of:

value_dim = hv.Dimension('Percentage', unit='%')

But can't find a way to have both values in the label.

Additionally, I tried to modify the hover tag. In my search to find ways to modify this I found ways to reference and display various attributes in the hover information (through the bokeh tooltips) but it does not seem like you can manipulate this information.

1 Answers

In this post two possible ways are explained how to achive the wanted result. Let's start with the example DataFrame and the necessary imports.

import holoviews as hv
from holoviews import opts, dim # only needed for 2. solution
import pandas as pd

data = {'A':['XX','XY','YY','XY','XX','XX'],
        'B':['RR','KK','KK','RR','RK','KK'],
        'values':[10,5,8,15,19,1],
       }

df = pd.DataFrame(data)

1. Option Use hv.Dimension(spec, **params), which gives you the opportunity to apply a formatter with the keyword value_format to a column name. This formatter is simple the combination of the value and the value in percent.

total = df.groupby('A', sort=False)['values'].sum().sum()

def fmt(x):
    return f'{x} ({round(x/total,2)}%)'

hv.Sankey(df, vdims = hv.Dimension('values', value_format=fmt))

2. Option Extend the DataFrame df by one column wich stores the labels, you want to use. This can be later reused inside the Sankey, with opts(labels=dim('labels')). To ckeck if the calculations are correct, you can turn show_values on, but this will cause a duplicate inside the labels. Therefor in the final solution show_values is set to False. This can be sometime tricky to find the correct order.

labels = []
for item in ['A', 'B']:
    grouper = df.groupby(item, sort=False)['values']
    total_sum = grouper.sum().sum()
    for name, group in grouper:
        _sum = group.sum()
        _percent = round(_sum/total_sum,2)
        labels.append(f'{name} - {_sum} ({_percent}%)')
df['labels'] = labels

hv.Sankey(df).opts(show_values=False, labels=dim('labels'))

The downside of this solution is, that we apply a groupby for both columns 'A' and 'B'. This is something holoviews will do, too. So this is not very efficient.

Output

Sanky with new Labels

Comment

Both solutions create nearly the same figure, except that the HoverTool is not equal.

Related