We have a large dataset of health records (1 row per patient) with several columns, each indicating whether or not the patient interacted with a particular type of healthcare provider (0=no, 1=yes). We are hoping to identify the combination of "yes"s (i.e., which providers were seen) for each patient.
The answers to this question get me a very long way toward my final goal, but I would like to find a way to assign slightly more human-readable names to the identified combinations of 0s and 1s.
The code below yields a toy dataset containing a factor (named "combo" here) with values consisting of 1s and 0s listed in the order in which they appear in the columns, separated by periods (e.g., 1.1.1.0.1.1).
df <- read.table(text =
"ID Pr1 Pr2 Pr3 Pr4 Pr5 Pr6
1 1 1 1 0 1 1
2 0 0 1 1 0 1
3 0 0 1 1 0 1
4 0 1 0 0 1 1
5 0 1 0 1 1 1
6 0 1 0 1 1 1
7 1 1 1 1 1 1
8 0 1 0 1 1 1
9 0 0 0 0 0 1
", header = TRUE)
combo <- do.call(interaction,c(df[-1],drop=TRUE))
df.new <- cbind(df, combo)
Because the real dataset has so many columns of 0/1 variables and potentially hundreds of observed combinations of 0s and 1s, these kinds of strings are going to be difficult to link back to the meaningful column names.
To make this connection a bit easier, what I would like to have is a new character or factor column with values that contain only the names of columns that have a value of 1, e.g., a combo value of 1.1.1.0.1.1 would yield a new value of "Pr1.Pr2.Pr3.Pr5.Pr6" and 0.0.0.0.0.1 would yield "Pr6". Even something like "Pr1.Pr2.Pr3.x.Pr5.Pr6" (or "x.x.x.x.x.Pr6") would be a bit easier to use than the original result.
Thanks for any assistance you can provide!