R find matches for more than 2 vectors

Viewed 728

I am working with a set of 5 excel columns A,B,C,D,E of words "Aaa","Aab"... and I want to find the exact matches in all the columns (in R).

A   B   C   D   E  
Aaa Aaa Baa Aaa Ass
Aab Ccc Aaa Baa Aaa
Ccc Abc Ccc Ccc Ccc
... ... ... ... ... 

I create a vector for each column.
For that I have try a for loop with if and grep function.

<pre>
    for(i in A_vector) {
          if(grep("i", B_vector))
              if(grep("i", C_vector))
                  if(grep("i", D_vector))
                      if(grep("i", E_vector))
                          print(i)
      }
<code>

(but I only obtain the words in the first vector A_vector).
At the end I would like to have a vector with the words "Aaa", "Bbb"... that match in the 5 columns. I do not need the position of each match within the vector, just the words that are common to all the vectors.

 Result
    [1] "Aaa"
    [2] "Ccc"
    [n]  ...

Thank you in advance!

3 Answers

You are asking to find common elements between each list, not just duplicates in general. Duplicates below are Aaa, Ccc, Ddd, and Xxx, but the only element duplicated across any is Xxx. intersect() will accomplish this, with some double lapply functions.

A = list("Aaa", "Aaa", "Ccc", "Ccc")
B = list("Ddd", "Ddd", "Ddd", "Eee")
C = list("Fff", "Ggg", "Hhh", "Iii", "Jjj")
D = list("Kkk", "Lll", "Mmm", "Nnn", "Xxx")
E = list("Ppp", "Qqq", "Rrr", "Xxx")
Mylist <- list(A, B, C, D, E)

dupes <- unlist(lapply(Mylist, function(x) lapply(Mylist, function(y) intersect(x,y))))

unique(dupes[duplicated(dupes)])

[1] "Xxx"

To see where the intersections are, this will tell you that your 4th list has 1 element in common with your 5th list:

sapply(seq_len(length(Mylist)), function(x) sapply(seq_len(length(Mylist)), function(y) length(intersect(unlist(Mylist[x]), unlist(Mylist[y])))))

     [,1] [,2] [,3] [,4] [,5]
[1,]    2    0    0    0    0
[2,]    0    2    0    0    0
[3,]    0    0    5    0    0
[4,]    0    0    0    5    1
[5,]    0    0    0    1    4

You could try something, though a little convoluted, using data.table:

library(data.table)

setDT(data)

data[, unlist(lapply(.SD, intersect, y = unique(A))), A][, .N, A][N == {ncol(dt) - 1}, A]

Here is the edited answer based on your explanation that you want to find all the matches between at least two of the columns:

 Mylist <-list(A=c("Aaa","Aab","Ccc","Ddd"), B=c("Aaa","Ccc","Abc","Abd"), C=c("Baa","Aaa","Ccc","Abb","Ddd"), D=c("Aaa","Baa","Ccc","CBB","Baa"),E=c("Ass","Aaa","Ccc","Gef"))
 CharVec <-unlist(Mylist)
 unique(CharVec[duplicated(CharVec)])
Related