Collapsing data to a single row based on ID

Viewed 69

I am tryng to collapse my data into a single row based on ID. The data consists of IDs (site locations) and the amount of species found in each site (SP1, SP2 etc). At the moment their are multiple IDs depending on how many individual species are found on that site and I want to collapse the data down so I just have a single row based on the site ID.

My data currently looks similar to this:

my_data <- data.frame(ID = c(1,2,3,3,3,4,4,4,4),
                      SP1 = c(10,15,0,0,0,0,0,0,10),
                      SP2 = c(0,0,5,0,0,8,0,0,0),
                      SP3 = c(0,0,0,20,0,0,4,0,0),
                      SP4 = c(0,0,0,0,0,0,0,12,0),
                      SP5 = c(0,0,0,0,6,0,0,0,0))

and I want to get it to look like this:

my_data2 <- data.frame(ID = c(1,2,3,4),
                       SP1 = c(10,15,0,10),
                       SP2 = c(0,0,5,8),
                       SP3 = c(0,0,20,4),
                       SP4 = c(0,0,0,12),
                       SP5 = c(0,0,6,0))

I think the best way to do it would be something along the lines of the code below but not sure on what to include in the summarise

my_data %>%
        group_by(ID)
        summarise()
1 Answers

It's not clear from your description what type of summary you want to calculate. sum, max, or possibly something else would produce the result you want. That said, the following would work. Here I use sum, but you could easily substitute another function:

library(tidyverse)

my_data2 <- my_data %>% 
  group_by(ID) %>% 
  summarize(
    across(everything(), sum)
  )

# A tibble: 4 × 6
     ID   SP1   SP2   SP3   SP4   SP5
  <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
1     1    10     0     0     0     0
2     2    15     0     0     0     0
3     3     0     5    20     0     6
4     4    10     8     4    12     0
Related