Find Positive Association and Negative Association Between 2 Categorical Variables

Viewed 46

I have different employees’ performance rating data( low, mixed, strong, significant, exceptional) in 1 column and certain commonly occurring words used in their performance reviews encoded as 0 and 1 (0 meaning the word is not present,1 meaning the word is present in their Perf reviews). For example have multiple columns like “leadership”, “excellent”,”lacking”… etc etc which are encoded as 0/1 for each employee example:

empID perf rating team leadership lacking excellent good
A123 low 1 0 1 1 0
C453 mixed 1 1 0 0 0
B335 strong 0 0 1 0 1
F976 significant 1 0 1 1 0
G257 exceptional 1 1 1 1 0

I need to find out which words are associated with a positive performance and which words are associated with negative performance? The output dataframe should be in the form of word, correlation-coefficient.

  1. I understand that since these are both qualitative variables, we cannot use Pearson's correlation coefficient method and we can use something like Cramer's v to find out the correlation coefficients. But I need the correlation coefficients to be between -1 and 1 rather than 0 to 1, to tell which words are positively associated with Performance rating and which words are negatively associated with Performance rating.
  2. If I encode the performance rating to 1,2,3,4,5 1 being low and 5 being exceptional, and since the presence of a word is also in the form of 0/1, can I still use Pearson's correlation coefficient to get positive association and negative association of words with performance rating? or is it a blunder?

For example, my output should be something like

word corr-coeff
team -0.02
leadership 0.712
lacking -0.8122
excellent 0.6172
good 0.5672
1 Answers

There are a few issues to bear in mind here.

  • you have repeated meaures on empID. That is, observations on one exmployee are likely to be more similar to each other than to observations on other employees. This means the observations are not independent and this needs to be accounted for.

  • the research question seems to warrant a regression model

  • I would consider using a multinomial logistic model with random intercepts for employee

  • this will provide estimates for the association between each commonly occuring word and the performance ratings, while accounting for the non-independence of observations within employees.

Related