Performing a logistic regression in R, several missing values I can't explain from the dataset

Viewed 179

I am performing a logistic regression on medical data in a data table with 370 rows. The dependent variable is the outcome of cirrosis (yes/no, or 0/1) (named fibrosis.score in the dataset), this is a factor with 2 levels 0/1.

In my dataset, the rows (one patient per row) have observations like: CT-scan: "no CT scan", "yes with cirrosis", "yes with steatosis", "yes without cirrosis". This variable is a factor with 4 levels (0-3). I have similar variables for US-scan and MR-scan, in addition to some clinical variables as a symptom yes/no, e.g. ascites 0/1.

So all my independent variables are factors with 2-4 levels.

When I try to perform the logistic regression, the results come back basically making me think there is something wrong with either the category of the variables, the reference levels or something with NAs.

    > model0 <- glm(Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites + ikterus + varicebloedning, family = binomial, data=d3)

Warning message: glm.fit: fitted probabilities numerically 0 or 1 occurred

summary(model0)

    Call:
    glm(formula = Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites + 
ikterus + varicebloedning, family = binomial, data = d3)

    Deviance Residuals: 
        Min       1Q   Median       3Q      Max  
    -1.4384  -0.2789  -0.2789  -0.2078   2.7736  

    Coefficients:
                       Estimate Std. Error z value Pr(>|z|)    
    (Intercept)         -3.8247     0.6608  -5.788 7.12e-09 ***
    Fibroscan1           2.6589     0.5432   4.895 9.83e-07 ***
    CT1                  1.4983     0.8796   1.704   0.0885 .  
    CT2                 21.5538  3200.8814   0.007   0.9946    
    CT3                  0.7466     0.8930   0.836   0.4031    
    MR1                -15.9709  6173.3230  -0.003   0.9979    
    MR2                 18.9707 10754.0130   0.002   0.9986    
    MR3                -17.6565  3611.3857  -0.005   0.9961    
    UL1                -16.3383  3131.5724  -0.005   0.9958    
    UL2                  4.4201     1.0597   4.171 3.03e-05 ***
    UL3                  0.5976     0.6391   0.935   0.3498    
    ascites1            17.3453  4572.5578   0.004   0.9970    
    ikterus1            20.1343 10754.0129   0.002   0.9985    
    varicebloedning1    -1.9442  6320.3723   0.000   0.9998    

Signif. codes: 0 ‘’ 0.001 ‘’ 0.01 ‘’ 0.05 ‘.’ 0.1 ‘ ’ 1

(Dispersion parameter for binomial family taken to be 1)

Null deviance: 218.43  on 245  degrees of freedom

Residual deviance: 102.72 on 232 degrees of freedom (124 observations deleted due to missingness) AIC: 130.72

Number of Fisher Scoring iterations: 18

  1. First of all - 124 observations are deleted due to missingness - however, there are no missing values in each of these independent variables in the dataset - is the glm code reliant on no missing values in the whole row (even though a lot of columns are not included in the glm)?

  2. Second - The p-values of almost all the independent variables are so "extreme" (i.e. 0.99), that i feel like there is something off with the reference values. The three last covariates in the code, should be significant according to our knowledge. However, the two significant covariates should definitely also be significant.

  3. I think some of the problem relies in the exclusion of the 124 missing values, I don't really understand how to come around this problem, making a whole new dataset with only the variables I am using in the regression?

Any explanation and tips would be greatly appreciated.

Update:

    > form <- Fibrosis.score ~ Fibroscan + CT + MR + UL + ascites + 
    ikterus + varicebloedning  
    > summary(d3[all.variables(form)])
    Error in all.variables(form) : could not find function 
    "all.variables"

Update 2: I am not sure I can provide an anonymized data set but I'll look into it! There are several NAs in the dataset as a whole, I have 370 obs of 21 variables. But the variables used in the regression do not contain NAs, this is why I am confused, they should be independent on the other variables?

Update 3: Problem solved. One of the dependent variables had 124 missing values, which I had overlooked. When this was omitted from the regression, the results become more what we expected.

Thank you all for input and help.

0 Answers
Related