I have the following dataframe:
df1 = pd.DataFrame(data={'1': ['a', 'd', 'g', 'j'],
'2': ['b', 'e', 'h', 'k'],
'3': ['c', 'f', 'i', 'l'],
'top_n': [1, 3, 2, 1]},
index=pd.Series(['ind1', 'ind2', 'ind3', 'ind4'], name='index'))
>>> df1
1 2 3 top_n
index
ind1 a b c 1
ind2 d e f 3
ind3 g h i 2
ind4 j k l 1
How would I only get the first N values for each row based on the top_n column?
>>> df1
1 2 3 top_n
index
ind1 a NaN NaN 1
ind2 d e f 3
ind3 g h NaN 2
ind4 j NaN NaN 1
In this example, ind3 has g and h because the top_n values is 2.