I'm trying to find a nice way to visualize the data from publicly available cBioPortal mutation data. I want to plot the co-occurrence of Protein Change (so basically, for each sample ID, does that specific sample have any other mutation also). See the image below:
I want to get this plotted as a heatmap (example below):
I've managed to get the data into the form of the first image above, but I am completely stuck as to how to go from there to the example heatmap.
I've looked into:
df.groupby(['Sample ID', 'Protein Change', 'Cancer Type Detailed']).count().unstack('Protein Change'))
which seems to go in the right direction, but not completely there.
Basically what I want is a heatmap with Protein Change on both axis, and a count of how many times those co-exist within a single sample.
Any help would be appreciated. Thanks!!


