How to prepare conjoint data in R?

Viewed 157

I have a question regarding the manipulation of my data in R. I am analyzing conjoint data and need to manipulate my choice variable to fit a given standard. Below is how my dataframe currently looks per respondent.

However to use it with the package 'ChoiceModelR' I need to change it so that it looks like in the second table. Currently the Choice variable is a binary variable indicating which alternative has been selected. In the required format the choice variable is always shown in the first row of a new question an refers to which alternative has been selected. When alternative 2 was selected in question 1, the choice variable will be 2 on the first row of question 1. If alternative 2 has been selected in question 2, the choice variable will be 1 on the first row of question 1. The second row of a question will always be 0 in this case.

The required format is given as the second table below.

Is there an easy way to code this in R?

My current data:

|   | ID | Question | Alternative | Choice | X_1 | X_2 | X_3 | X_4 | X_5 | X_6 | X_7 |
|---|----|----------|-------------|--------|-----|-----|-----|-----|-----|-----|-----|
|   | 1  | 1        | 1           | 0      | 2   | 2   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 1        | 2           | 1      | 2   | 2   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 2        | 1           | 1      | 1   | 1   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 2        | 2           | 0      | 2   | 1   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 3        | 1           | 0      | 2   | 1   | 2   | 1   | 1   | 2   | 1   |
|   | 1  | 3        | 2           | 1      | 1   | 2   | 2   | 2   | 1   | 2   | 2   |
|   | 1  | 4        | 1           | 0      | 1   | 1   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 4        | 2           | 1      | 1   | 2   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 5        | 1           | 1      | 2   | 1   | 2   | 2   | 1   | 2   | 1   |
|   | 1  | 5        | 2           | 0      | 2   | 1   | 1   | 1   | 2   | 1   | 1   |

How it has to look:

|   | ID | Question | Alternative | Choice | X_1 | X_2 | X_3 | X_4 | X_5 | X_6 | X_7 |
|---|----|----------|-------------|--------|-----|-----|-----|-----|-----|-----|-----|
|   | 1  | 1        | 1           | 2      | 2   | 2   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 1        | 2           | 0      | 2   | 2   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 2        | 1           | 1      | 1   | 1   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 2        | 2           | 0      | 2   | 1   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 3        | 1           | 2      | 2   | 1   | 2   | 1   | 1   | 2   | 1   |
|   | 1  | 3        | 2           | 0      | 1   | 2   | 2   | 2   | 1   | 2   | 2   |
|   | 1  | 4        | 1           | 2      | 1   | 1   | 1   | 1   | 2   | 1   | 1   |
|   | 1  | 4        | 2           | 0      | 1   | 2   | 1   | 1   | 2   | 1   | 2   |
|   | 1  | 5        | 1           | 1      | 2   | 1   | 2   | 2   | 1   | 2   | 1   |
|   | 1  | 5        | 2           | 0      | 2   | 1   | 1   | 1   | 2   | 1   | 1   |

UPDATE 14th of JUNE 2020

In case anyone runs into the same problem I have found a way to format the data correctly. The code I have used is displayed below.

choice <- rep(0, nrow(your_df)) #your_df is your dataframe, creates a vector of 0's that is the length of your_df. 
choice[your_df[,"alternative"]==1] <- your_df[your_df[,"choice"]==1,"alternative"] # formats the data in the correct way
new_df <- cbind(your_df, choice) #merges your_df and choice
new+df = subset(new_df, select = -c(selected)) # remove the original selected column
2 Answers

This seems to transform the data in the way that you'd like, though it's hard to understand your exact conditions:

library(dplyr)

df %>% 
  group_by(Question) %>% 
  mutate(Choice = 
           case_when(
             Question %in% c(1, 3, 4) & Alternative == 2 ~ 2,
             Question %in% c(2,5) & Alternative == 2 ~ 1
           ),
         Choice = lead(Choice)) %>% 
  replace(is.na(.), 0)

Gives us:

      ID Question Alternative Choice   X_1   X_2   X_3   X_4   X_5   X_6   X_7
   <dbl>    <dbl>       <dbl>  <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl> <dbl>
 1     1        1           1      2     2     2     1     1     2     1     1
 2     1        1           2      0     2     2     1     1     2     1     2
 3     1        2           1      1     1     1     1     1     2     1     1
 4     1        2           2      0     2     1     1     1     2     1     2
 5     1        3           1      2     2     1     2     1     1     2     1
 6     1        3           2      0     1     2     2     2     1     2     2
 7     1        4           1      2     1     1     1     1     2     1     1
 8     1        4           2      0     1     2     1     1     2     1     2
 9     1        5           1      1     2     1     2     2     1     2     1
10     1        5           2      0     2     1     1     1     2     1     1

Data:

df <- structure(list(ID = c(1, 1, 1, 1, 1, 1, 1, 1, 1, 1), Question = c(1, 
1, 2, 2, 3, 3, 4, 4, 5, 5), Alternative = c(1, 2, 1, 2, 1, 2, 
1, 2, 1, 2), Choice = c(0, 1, 1, 0, 0, 1, 0, 1, 1, 0), X_1 = c(2, 
2, 1, 2, 2, 1, 1, 1, 2, 2), X_2 = c(2, 2, 1, 1, 1, 2, 1, 2, 1, 
1), X_3 = c(1, 1, 1, 1, 2, 2, 1, 1, 2, 1), X_4 = c(1, 1, 1, 1, 
1, 2, 1, 1, 2, 1), X_5 = c(2, 2, 2, 2, 1, 1, 2, 2, 1, 2), X_6 = c(1, 
1, 1, 1, 2, 2, 1, 1, 2, 1), X_7 = c(1, 2, 1, 2, 1, 2, 1, 2, 1, 
1)), row.names = c(NA, -10L), class = c("tbl_df", "tbl", "data.frame"
))

Does this work for you?

  • df is the original choice data

 df <- data.frame(
      ID = c(1, 1,  1,  1,  1,  1,  1,  1,  1,  1),
      Question = c(1,   1,  2,  2,  3,  3,  4,  4,  5,  5),
      Alternative = c(1,    2,  1,  2,  1,  2,  1,  2,  1,  2),
      Choice =  c(0,    1,  1,  0,  0,  1,  0,  1,  1,  0),
      X_1  =    c(2,    2,  1,  2,  2,  1,  1,  1,  2,  2),
      X_2 = c(2,    2,  1,  1,  1,  2,  1,  2,  1,  1),
      X_3 = c(1,    1,  1,  1,  2,  2,  1,  1,  2,  1),
      X_4 = c(1,    1,  1,  1,  1,  2,  1,  1,  2,  1),
      X_5 = c(2,    2,  2,  2,  1,  1,  2,  2,  1,  2),
      X_6 = c(1,    1,  1,  1,  2,  2,  1,  1,  2,  1),
      X_7 = c(1,    2,  1,  2,  1,  2,  1,  2,  1,  1)
      )

df

   ID   Question Alternative Choice X_1 X_2 X_3 X_4 X_5 X_6 X_7
1   1        1           1      0   2   2   1   1   2   1   1
2   1        1           2      1   2   2   1   1   2   1   2
3   1        2           1      1   1   1   1   1   2   1   1
4   1        2           2      0   2   1   1   1   2   1   2
5   1        3           1      0   2   1   2   1   1   2   1
6   1        3           2      1   1   2   2   2   1   2   2
7   1        4           1      0   1   1   1   1   2   1   1
8   1        4           2      1   1   2   1   1   2   1   2
9   1        5           1      1   2   1   2   2   1   2   1
10  1        5           2      0   2   1   1   1   2   1   1

Create a new data set df2 with a new variable DepVar that recodes the Choice variable. (Or, you can ignore the df2 part, just modify the df itself)

df2 <- df %>% mutate(DepVar = ifelse(Choice==1, Alternative, 0)) %>%
           arrange(ID, Question, -DepVar)

df2

   ID   Question Alternative Choice X_1 X_2 X_3 X_4 X_5 X_6 X_7 DepVar
1   1        1           2      1   2   2   1   1   2   1   2      2
2   1        1           1      0   2   2   1   1   2   1   1      0
3   1        2           1      1   1   1   1   1   2   1   1      1
4   1        2           2      0   2   1   1   1   2   1   2      0
5   1        3           2      1   1   2   2   2   1   2   2      2
6   1        3           1      0   2   1   2   1   1   2   1      0
7   1        4           2      1   1   2   1   1   2   1   2      2
8   1        4           1      0   1   1   1   1   2   1   1      0
9   1        5           1      1   2   1   2   2   1   2   1      1
10  1        5           2      0   2   1   1   1   2   1   1      0
Related