I have a very large pairwise distance matrix in R. I'd like to code cell in the matrix based on whether the row/column names are the same or different.
On a smaller scale, the row/column names would be:
individuals <- c("apple", "pear", "apple", "cranberry", "peach", "apple")
I would like a matrix with 1 for each comparison involving apple, except for comparisons of apple to apple. That would look like:
[,1] [,2] [,3] [,4] [,5] [,6]
[1,] "0" "1" "1" "1" "1" "1"
[2,] "1" "0" "1" "0" "0" "1"
[3,] "1" "1" "0" "1" "1" "1"
[4,] "1" "0" "1" "0" "0" "1"
[5,] "1" "0" "1" "0" "0" "1"
[6,] "1" "1" "1" "1" "1" "0"
I know I can achieve this by doing:
final.matrix <- matrix(nrow= length(individuals), ncol = length(individuals))
final.matrix[grep("apples", individuals),] <- 1
final.matrix[,grep("apples", individuals)] <- 1
diag(final.matrix) <- 0
final.matrix[is.na(final.matrix)] <- 0
But there's gotta be a cleaner/simpler way. What am I missing?
Additionally, this doesn't work when the row/column names are a tibble, which is how they are in reality. Suggestions for a solution that works with tibbles?
tibble_inds <- as_tibble(individuals)
grep("apple", tibble_inds)
# 1