I have a dataframe that has a few million rows. I need to calculate the sum of each row from a particular column index up until the last column. The column index for each row is unique. An example of this, with the desired output, will be:
import pandas as pd
df = pd.DataFrame({'col1': [1, 2, 2, 5, None, 4],
'col2': [4, 2, 4, 2, None, 1],
'col3': [6, 3, 8, 6, None, 4],
'col4': [9, 8, 9, 3, None, 5],
'col5': [1, 3, 0, 1, None, 7],
})
df_ind = pd.DataFrame({'ind': [1, 0, 3, 4, 3, 5]})
for i in df.index.to_list():
df.loc[i, "total"] = df.loc[i][(df_ind.loc[i, "ind"]).astype(int):].sum()
print(df)
>>
col1 col2 col3 col4 col5 total
0 1.0 4.0 6.0 9.0 1.0 20.0
1 2.0 2.0 3.0 8.0 3.0 18.0
2 2.0 4.0 8.0 9.0 0.0 9.0
3 5.0 2.0 6.0 3.0 1.0 1.0
4 NaN NaN NaN NaN NaN 0.0
5 4.0 1.0 4.0 5.0 7.0 0.0
How can I achieve this efficiently with pandas without using a for loop. Thanks