Dear stackoverflow community,
This is my first question. I gave my best to make everything reproducible, so please show mercy if there are still things to improve ;-)
About the problem
I am a vet epidemiologist and analysing data from different diagnostic tests (macroscopical lesions, immunohistochemistry staining, PCR, ...). We aim to investigate whether those tests correlate with each other (e.g. "Animals with macroscopical lesions are also positive in the laboratory tests.").
The different diagnostic tests were used on the same foot and/or biopsy: the foot was scored macroscopically, a biopsy was taken and multiple laboratory tests were done from the same biopsy. As the data is paired (multiple samples from the same foot/biopsy), I decided to use McNemar's test.
However, when analysing the data, I have encountered some problems that confuse me. Here are a few scenarios:
Examples
Scenario A
df_A = data.frame(test1 = c(rep(0, 20+11), rep(1, 36+49)),
test2 = c(rep(0, 20), rep(1, 11), rep(0, 36), rep(1, 49)))
table(df_A$test1, df_A$test2)
ggplot(df_A, aes(x = factor(test1), fill = factor(test2))) + geom_bar(position = "fill") +
geom_text(stat="count", aes(label=..count..), position=position_fill(vjust=0.5))
mcnemar.test(df_A$test1, df_A$test2)
This makes sense to me. In the plot, it's easy to see that those testing positive in test1 were also more likely to be positve in test2. (The odds ratio is 2.47.) McNemar gives me "p-value = 0.0004639". So far, so good.
Scenario B
df_B = data.frame(test1 = c(rep(0, 9+24), rep(1, 21+65)),
test2 = c(rep(0, 9), rep(1, 24), rep(0, 21), rep(1, 65)))
table(df_B$test1, df_B$test2)
ggplot(df_B, aes(x = factor(test1), fill = factor(test2))) + geom_bar(position = "fill") +
geom_text(stat="count", aes(label=..count..), position=position_fill(vjust=0.5))
mcnemar.test(df_B$test1, df_B$test2)
This still comes out as I expect. In the plot, I barely see a difference. The odds ratio is 1.16. McNemar gives me "p-value = 0.7656". Still: So far, so good.
Scenario C
df_C = data.frame(test1 = c(rep(0, 13+18), rep(1, 18+71)),
test2 = c(rep(0, 13), rep(1,18), rep(0, 18), rep(1, 71)))
table(df_C$test1, df_C$test2)
ggplot(df_C, aes(x = factor(test1), fill = factor(test2))) + geom_bar(position = "fill") +
geom_text(stat="count", aes(label=..count..), position=position_fill(vjust=0.5))
mcnemar.test(df_C$test1, df_C$test2)
Now, I start to be confused: In the plot, I again see quite a difference between the tests. The odds ratio is 2.85. However, McNemar gives me "p-value = 1".
Scenario D
df_D = data.frame(test1 = c(rep(0, 17+16), rep(1, 39+44)),
test2 = c(rep(0, 17), rep(1,16), rep(0, 39), rep(1, 44)))
table(df_D$test1, df_D$test2)
ggplot(df_D, aes(x = factor(test1), fill = factor(test2))) + geom_bar(position = "fill") +
geom_text(stat="count", aes(label=..count..), position=position_fill(vjust=0.5))
mcnemar.test(df_D$test1, df_D$test2)
Basically the "opposite" of scenario C: I don't see any difference between the two groups in the plot. The odds ratio is 1.20. Still, McNemar gives me "p-value = 0.003012".
What I did so far
Reading about McNemar's test in a little more detail, I figured out that basically only the two "discordant" fields matter. In scenario B, 21 and 24 are quite close. In scenario C, 18 and 18 are even equal. This explains why McNemar delivers a non-significant result. On the other hand, in scenario D, 16 and 39 are quite different, hence the significant result, even though the proportions are quite the same.
I also found this other question, in which this exact "behaviour" of the test is explained. However, the author seems to have the same interpretation difficulty of "comparing two groups" but not receiving the expected results from McNemar. However, this ("second") part of the question was never answered.
On Wikipedia, it states: "We also have to find out where these two tests disagree with each other. This is precisely the basis of McNemar's test." As far as I understand, this is exactly what I am trying to do. But the results are not really what I expected.
Conclusion - question
So, I am wondering now: Is my understanding of McNemar's test somehow wrong? Is it not appropriate for a statement like "Patients positive in one test are also likely to be positive in the other."? If so, which test could I use as an alternative?
Many thanks for your suggestions already in advance. Best, Brian