I have a datatable / dataframe with some columns of logical vectors - each logical vector represents the presence or absence of a species in a weekly survey. I want to identify those rows where the species is present over a 21 day period or more - the period that the species isn't recorded in between is irrelevant).
This means I need to identify all those rows where a 'TRUE' is present in at least 2 columns (only those columns with "detected" in the colname in example below) that are 3 or more columns apart (because each column represents a weekly survey).
So if e.g. V1 and V4 are TRUE, my 'result' column is TRUE. If only V6,V7,V8 are TRUE, my result column is FALSE. If V2 and V9 are TRUE, my result is TRUE etc.
This is part of a bigger simulation, and the number of columns in the dt (and the number of weekly surveys) varies depending on other simulation parameters, so it isn't possible to do it using column indexing. But the columns of interest can all be given the same suffix ('detected') as in the example code below.
There are also other TRUE/FALSE columns in the dt (that aren't suffixed with 'detected').
example data:
library(data.table)
set.seed(123)
#first column
colref <- 1:20
#number of columns will vary in the dt, so here we generate a random number
n<-sample(5:10,1)
# function to create some logical vectors for the dt
create.col<-function(n){replicate(n,sample(c(TRUE,FALSE),20,replace=TRUE,prob = c(0.2,0.8)),simplify=FALSE)}
# create the dt
dt<-setDT(create.col(n))
# affix "detected" to columns of interest
colnames(dt) <- paste("detected", colnames(dt), sep = "_")
dt<-data.table(colref,dt)
dt
# need to create logical vector 'result' column identifying rows where TRUE is present in columns separated by three or more positions in the dt, with the 'detected' suffix