Rowwise mean on selected columns in R

Viewed 5578

Let's illustrate the problem on the famous iris dataset. I need to apply the selected function by rows but only on selected columns. Example goes as follows:

library(tidyverse)

iris %>%
  mutate_at(.funs = scale, .vars = vars(-c(Species))) %>%
  rowwise() %>% 
  mutate(my_mean=mean(c(Sepal.Length, Sepal.Width, Petal.Length, Petal.Width)))

So, first I scale all variables, excluding Species and then compute mean rowwise over all four numeric variables. However, in the real dataset I have 100+ numeric variables and I wonder how to convince R to automatically include all variables excluding selected one (e.g., Species in the given example). I go through the solutions on SO (e.g., this), but all examples explicitly refer to column names. Any pointers are greatly welcome.

EDIT: after some munging here is my solution:

iris %>%
  as_tibble() %>% 
  mutate_at(.funs = scale, .vars = vars(-c(Species))) %>% 
  transmute(Species, row_mean = rowMeans(select(., -Species)))
1 Answers

I'm not sure I got exactly what the problem is, but here are a few alternative dplyr solutions which will give you the mean of all columns except the selected one:

iris %>%
    select(-Species) %>%
    mutate(Means = rowMeans(.))

iris %>%
    mutate(Means = rowMeans(.[,1:4]))

iris %>%
    mutate(Means = rowMeans(.[,-5]))

The first is the only one that eliminates the selected column from the return. Hope one of them helps you.

Related