have the following data frame lets call it df, with the following observations
| id | b | c | f | e_7 | ic_107 | d | g | j |
|---|---|---|---|---|---|---|---|---|
| 1 | 23 | 3 | 66 | 97 | 8 | 5 | 7 | 0 |
| 2 | 1 | 1 | 5 | 7 | NA | NA | NA | NA |
| 3 | NA | 2 | 79 | 5 | 5 | 4 | 9 | 0 |
| 4 | 0 | 2 | 32 | 1 | 6 | 6 | 1 | 0 |
| 5 | 36 | 6 | 9 | 49 | 9 | NA | NA | NA |
| 6 | 0 | 2 | 32 | 1 | 6 | 7 | 8 | 9 |
| 7 | 36 | NA | NA | 49 | 9 | 0 | 0 | 1 |
I want to retain only the records which do not have NA in many, but not all, columns. Let's say, column b, c, d, g, and j.
I am currently using filter with pipes, but I would like to avoid coding like:
df_new <- df %>%
filter(!is.na(b))%>%
filter(!is.na(c))%>%
filter(!is.na(d))%>%
filter(!is.na(g))%>%
filter(!is.na(j))
Is there an easier way to write the code?
In this example, I have 5 columns for the filtering condition. In my real dataset, I have 17. Therefore, I would like to avoid the coding above.
Also, instead of simple column names a, b, c, d..., the columns of my real dataset have long names, such as lighteningdate, depression,anxiety..., so I'd like to use a vector of column numbers (c(3:9, 13:21))rather than a list of column names in the coding.