I would like to be able to use dplyr to average rows in which there are identical values in ANY n or more numerical columns, and an identical value in the a column.
If:
n <- 3
and
df <- data.frame(a = c("one", "one", "one", "one", "three"),
b = c(1,1,1,2,3),
c = c(2,2,2,7,12),
d = c(6,6,7,8,10),
e = c(1,4,1,3,4))
then I would like the first three rows to be averaged (because 3 out of 4 numerical values are identical between them, and the value in a is also identical). I would NOT want row four to be included in the average, because although the value in a is identical, it has no identical numerical values.
Before:
a b c d e
[1] one 1 2 6 1
[2] one 1 2 6 4
[3] one 1 2 7 1
[4] one 2 7 8 3
[5] four 3 12 10 4
After:
a b c d e
[1] one 1 2 6.3 2
[2] one 2 7 8 3
[3] four 3 12 10 4
My data frame is much bigger in real life and contains plenty of other columns.
EDIT:
Rows [1] and [2] have 3 identical values (in columns b, c and d. Rows [1] and [3] have 3 identical values (in columns b, c and e. This is why I want them averaged.
