I have prevalence data by non-exclusive categories/classifications. (e.g., a story could be 'amazing', 'boring', 'charming', 'dark', or any combination of the four.) Illustrative:
library(data.table)
set.seed(0)
results = as.data.table( expand.grid( rep( list(0:1) , 4 ) ) )
names(results) = c('a', 'b', 'c', 'd')
results$prevalence = runif( n = 16 )
results$prevalence = results$prevalence/sum(results$prevalence)
I'd like to be able to answer the question(s):
- (trivial) What is the population coverage that is not in any category (
a = b = c = d = 0)? - What is the one category that covers the largest percent of the population?
- What are the two categories that cover the largest percent of the population?
- ... and so on...
Effectively, I'd like to create a quasi-CDF where:
- I know that for data in the none category (i.e.,
a = b = c = d = 0) I cover 10% of the population. - I know that for data in either one or no categories, I can cover 21% of the population by limiting myself to category
c.
That is:
results[ ( a == 0 & b == 0 & d == 0 ) & rowSums( results[ , -'prevalence' ] ) <= 1 , sum(prevalence) ]
- I know that for data in either two, one, or no categories, I can cover 36% of the population by limiting myself to categories
bandc.
That is:
results[ ( a == 0 & d == 0 ) & rowSums( results[ , -'prevalence' ] ) <= 2 , sum(prevalence) ]
- I know that for data in either three, two, one, or no categories, I can cover 59% of the population by limiting myself to categories
a,b, andc.
That is:
results[ ( d == 0 ) & rowSums( results[ , -'prevalence' ] ) <= 3 , sum(prevalence) ]
- And, trivially, I know that for data in either four, three, two, one, or no categories, I can cover 100% of the population by limiting myself to each of the four categories (
a,b,c,d).
In this limited example, I just checked all possible categories to find the largest prevalence by grouping of allowable non-zero categories (actually, as you see by my code snippets, I was doing the inverse and finding prevalence by grouping categories that were restricted to zero).
How can I do this in a data.table way so that I don't have to brute force through the many combinations of dummy variables (columns) in my real summary data set?
I have suspicions that it might involve some clever use of .EACHI or lapply that I haven't been able to think of.