Which is the preferred way to filter dataset in Python?

Viewed 39

Filter dataset 02 telecom_usage.csv with multiple condition

  • the value of CustomerCareCalls starting with even digits (2, 4, 6 or 8)
  • the value of UnansweredCalls > BlockedCalls

First way:

df_2 = pd.read_csv('02 telecom_usage.csv')
df_3 = df_2[(df_2['CustomerCareCalls'].str.contains('^2|^4|^6|^8')&
                 df_2['UnansweredCalls']>df_2['BlockedCalls'])]

second way:

df_4 = pd.read_csv('02 telecom_usage.csv')
df_5 = df_4[df_4['CustomerCareCalls'].str.contains('^2|^4|^6|^8')]
    df_6 = df_5[df_5['UnansweredCalls']>df_5['BlockedCalls']]]

I've tried to run both queries with same dataset, but there are different results in both

the first filter results as many as 160 data from 5000 data, while the results of the second filter are 558 data from 5000 data

after I checked the results manually via excel, the results both met the request for multiple conditions.

I want to know, why the result is difference

0 Answers
Related