I am performing a logistic regression on medical data in a data table with 370 rows. The dependent variable is the outcome of cirrosis (yes/no, or 0/1) (named fibrosis.score in the dataset), this is a factor with 2 levels 0/1.
In my dataset, the rows (one patient per row) have observations like: CT-scan: "no CT scan", "yes with cirrosis", "yes with steatosis", "yes without cirrosis". This variable is a factor with 4 levels (0-3). I have similar variables for US-scan and MR-scan, in addition to some clinical variables as a symptom yes/no, e.g. ascites 0/1.
So all my independent variables are factors with 2-4 levels.
When I try to perform the logistic regression, the results come back basically making me think there is something wrong with either the category of the variables, the reference levels or something with NAs.
> model0 <- glm(Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites + ikterus + varicebloedning, family = binomial, data=d3)
Warning message: glm.fit: fitted probabilities numerically 0 or 1 occurred
summary(model0)
Call:
glm(formula = Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites +
ikterus + varicebloedning, family = binomial, data = d3)
Deviance Residuals:
Min 1Q Median 3Q Max
-1.4384 -0.2789 -0.2789 -0.2078 2.7736
Coefficients:
Estimate Std. Error z value Pr(>|z|)
(Intercept) -3.8247 0.6608 -5.788 7.12e-09 ***
Fibroscan1 2.6589 0.5432 4.895 9.83e-07 ***
CT1 1.4983 0.8796 1.704 0.0885 .
CT2 21.5538 3200.8814 0.007 0.9946
CT3 0.7466 0.8930 0.836 0.4031
MR1 -15.9709 6173.3230 -0.003 0.9979
MR2 18.9707 10754.0130 0.002 0.9986
MR3 -17.6565 3611.3857 -0.005 0.9961
UL1 -16.3383 3131.5724 -0.005 0.9958
UL2 4.4201 1.0597 4.171 3.03e-05 ***
UL3 0.5976 0.6391 0.935 0.3498
ascites1 17.3453 4572.5578 0.004 0.9970
ikterus1 20.1343 10754.0129 0.002 0.9985
varicebloedning1 -1.9442 6320.3723 0.000 0.9998
Signif. codes: 0 ‘’ 0.001 ‘’ 0.01 ‘’ 0.05 ‘.’ 0.1 ‘ ’ 1
(Dispersion parameter for binomial family taken to be 1)
Null deviance: 218.43 on 245 degrees of freedom
Residual deviance: 102.72 on 232 degrees of freedom (124 observations deleted due to missingness) AIC: 130.72
Number of Fisher Scoring iterations: 18
First of all - 124 observations are deleted due to missingness - however, there are no missing values in each of these independent variables in the dataset - is the glm code reliant on no missing values in the whole row (even though a lot of columns are not included in the glm)?
Second - The p-values of almost all the independent variables are so "extreme" (i.e. 0.99), that i feel like there is something off with the reference values. The three last covariates in the code, should be significant according to our knowledge. However, the two significant covariates should definitely also be significant.
I think some of the problem relies in the exclusion of the 124 missing values, I don't really understand how to come around this problem, making a whole new dataset with only the variables I am using in the regression?
Any explanation and tips would be greatly appreciated.
Update:
> form <- Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites +
ikterus + varicebloedning
> summary(d3[all.variables(form)])
Error in all.variables(form) : could not find function
"all.variables"
Update 2: I am not sure I can provide an anonymized data set but I'll look into it! There are several NAs in the dataset as a whole, I have 370 obs of 21 variables. But the variables used in the regression do not contain NAs, this is why I am confused, they should be independent on the other variables?
Update 3: Problem solved. One of the dependent variables had 124 missing values, which I had overlooked. When this was omitted from the regression, the results become more what we expected.
Thank you all for input and help.