I have an igraph network graph with 103,887 nodes and 4,795,466 ties.
This can be structured as an edgelist in a data.table with almost 9 million rows.
I can find the common neighbors in this network, following @chinsoon12's answer here. See the example below.
This works beautifully for smaller networks, but runs into problems in my use-case because the merge results in more than 2^31 rows.
Questions:
- Are there efficient alternatives on how to deal with this?
- Can I split the data and do the computation in steps? The results will be used to query about common neighbors.
Example - modified from @chinsoon12's answer:
library(data.table)
library(igraph)
set.seed(1234)
g <- random.graph.game(10, p=0.10)
adjSM <- as(get.adjacency(g), "dgTMatrix")
adjDT <- data.table(V1=adjSM@i+1, V2=adjSM@j+1)
res <- adjDT[adjDT, nomatch=0, on="V2", allow.cartesian=TRUE
][V1 < i.V1, .(Neighbours=paste(V2, collapse=",")),
by=c("V1","i.V1")][order(V1)]
res
V1 i.V1 Neighbours
1: 4 5 8
2: 4 10 8
3: 5 10 8