How to use slice in dplyr to keep the rows with NA values in R

Viewed 1360

I have the following dataset, and I want to know the min word for each group, and if there is no min word (it is NA), I still want to display it

df=data.frame(
  key=c("A","A","B","B","C"),
  word=c(1,2,3,5,NA))

df%>%group_by(key)%>%slice(which.min(word))

This excludes key=C, word=NA which I would want:

df_out=data.frame(
  key=c("A","B","C"),
  word=c(1,3,NA))
4 Answers

We can create a logical condition with is.na in filter and return the NA rows as well after doing the grouping by 'key'

library(dplyr)
df %>%
     group_by(key) %>% 
     filter(word == min(word)|is.na(word))

Or using slice. We don't need any if/else condition

df %>%
    group_by(key) %>% 
    slice(which(word ==min(word)|is.na(word)))
# A tibble: 3 x 2
# Groups:   key [3]
#  key    word
#  <chr> <dbl>
#1 A         1
#2 B         3
#3 C        NA

Or more compactly

df %>%
    group_by(key) %>% 
    slice(match(min(word), word))
# A tibble: 3 x 2
# Groups:   key [3]
#  key    word
#  <chr> <dbl>
#1 A         1
#2 B         3
#3 C        NA

NOTE: Using match returns the index of the first match.


which.min removes the NA

which.min(c(NA, 1, 3))
#[1] 2

We can check the condition with if, If all the word in a group is NA we return the first row or else return the minimum row.

library(dplyr)

df %>%
  group_by(key)%>%
  slice(if(all(is.na(word))) 1L else which.min(word))

# key    word
#  <chr> <dbl>
#1 A         1
#2 B         3
#3 C        NA

Another option is to arrange the data by word and select the 1st row in each group.

df %>% arrange(key, word) %>% group_by(key) %>% slice(1L)

You can also arrange by word and use distinct from dplyr to get the desired output.

library(dplyr)
df %>% 
    arrange(word) %>% 
    distinct(key, .keep_all = TRUE)
#  key word
#1   A    1
#2   B    3
#3   C   NA

You can create a modified slice-function using the tidyverse-package, which returns NA's:

slice_uneven = function(.data, .idx) {
  .data_ = .data %>% add_row() # Add an extra row
  .idx_ = .idx %>% c(NA) %>% replace_na(nrow(.data_)) # Replace NA with index of the extra row
  .data_[.idx_,] %>% head(-1) %>% remove_rownames() %>% return() # Subset, remove extra row, and reset rownames before returning data
}

slice_uneven(cars, c(1, 2, 3, NA, NA, 3, 2))
Related