Quick sum of all rows that fill a condition in DataFrame

Viewed 72

I have a pandas dataframe that looks something like this:

df = pd.DataFrame(np.array([[1,1, 0], [5, 1, 4], [7, 8, 9]]),columns=['a','b','c'])

   a  b  c
0  1  1  0
1  5  1  4
2  7  8  9

I want to find the first column in which the majority of elements in that column are equal to 1.0.

I currently have the following code, which works, but in practice, my dataframes usually have thousands of columns and this code is in a performance critical part of my application, so I wanted to know if there is a way to do this faster.

for col in df.columns:
    amount_votes = len(df[df[col] == 1.0])
    if amount_votes > len(df) / 2:
       return col

In this case, the code should return 'b', since that is the first column in which the majority of elements are equal to 1.0

2 Answers

Try:

print((df.eq(1).sum() > len(df) // 2).idxmax())

Prints:

b

Find columns with more than half of values equal to 1.0

cols = df.eq(1.0).sum().gt(len(df)/2)

Get first one:

cols[cols].head(1)
Related