randomForest missing independent variables in prediction

Viewed 47

I am working with a 14k row dataset with the following independent variable which have been mutated to factors. I am using randomForest to predict the future fall basedWhen I run the RF equation, the output only generates predictions for 2 of the 3 recent_fall variables. I have tried to change the ntree, mtry, and nodesize, but I am unable to generate a 3 variable prediction. I am still very new to R and data analysis so any suggestions/tips are greatly appreciated. Thanks!

Original dataset (training set)

full_id <- c("10019CNI", "10021YGW", "10043XFI", "10069AGJ", "10085CEW", "10093GTB", "10107OSZ", "10154ETO", "10174HKD", "10245RUI", "10270WID", "10292RMS", "10334TXS", "10367NMM", "10480SUN", "10482VXQ", "10489EHV", "10496MIO", "10524KKX", "10619WSC")

recent_fall <- c("no fall", "minor fall", "major fall", "major fall", "no fall", "no fall", "major fall", "minor fall", "major fall", "no fall", "minor fall", "minor fall", "major fall", "minor fall", "minor fall", "no fall", "no fall", "major fall", "major fall", "no fall")

er_visit <- c("hospitalized", "no visit", "no visit", "hospitalized", "no visit", "no visit", "no visit", "no visit", "no visit", "no visit", "no visit", "hospitalized", "hospitalized", "hospitalized", "hospitalized", "no visit", "no visit", "no visit", "hospitalized", "no visit")

age <- c(75, 89, 90, 90, 78, 80, 83, 91, 93, 72, 79, 82, 82, 90, 86, 70, 71, 93, 85, 95)

gender <- c("male", "male", "male", "female", "male", "male", "female", "female", "female", "male", "male", "male", "male", "male", "male", "female", "male", "male", "female", "female")

dx <- c("incontinence", "parkinsons", "stroke", "blind", "diabetes", "diabetes", "heart failure", "stroke", "heart attack", "uc", "incontinence", "cancer", "heart attack", "chf", "uc", "heart failure", "vertigo", "cancer", "alzheimers", "asthma")

RF test dataset (test set)

age_test <- c(74, 78, 79, 93, 87, 94, 86, 80, 90, 83, 82, 86, 74, 80, 78, 86, 80, 93, 72, 84)

gender_test <- c("female", "female", "male", "male", "female", "female", "male", "male", "female", "male", "male", "female", "female", "female", "male", "male", "male", "female", "female", "female")

dx_test <- c("asthma", "heart failure", "chf", "hearing loss", "incontinence", "diabetes", "chf", "copd", "ptsd", "diabetes", "diabetes", "stroke", "dementia", "copd", "vertigo", "none", "diabetes", "none", "cancer", "heart attack")

Dataframe structure (training set)

'data.frame':   14712 obs. of  9 variables:
 $ full_id         : chr  "10000JTS" "10001NVH" "10002TJZ" "10003NQJ" ...
 $ birth_month_year: chr  "Nov-45" "Aug-33" "May-43" "Nov-41" ...
 $ age             : int  76 88 78 80 78 91 73 76 78 69 ...
 $ gender          : Factor w/ 2 levels "female","male": 1 1 2 1 1 2 1 1 2 2 ...
 $ dx              : Factor w/ 22 levels "alzheimers","amputee",..: 14 17 19 5 10 10 13 19 12 21 ...
 $ recent_fall     : Factor w/ 3 levels "major fall","minor fall",..: 3 3 3 1 3 2 3 3 2 2 ...
 $ er_visit        : Factor w/ 2 levels "hospitalized",..: 2 2 2 1 2 1 2 2 2 1 ...
 $ id              : int  1 2 3 4 5 6 7 8 9 10 ...

Confusion matrix

randomForest(recent_fall~.,data=pt_data_new, ntree=500, mtry = 3)
Call:
 randomForest(formula = recent_fall ~ ., data = pt_data_new, ntree = 500,      mtry = 3) 
               Type of random forest: classification
                     Number of trees: 500
No. of variables tried at each split: 3

        OOB estimate of  error rate: 39.66%
Confusion matrix:
           major fall minor fall no fall class.error
major fall        359       1524     883   0.8702097
minor fall        567       2838    1830   0.4578797
no fall           159        872    5680   0.1536284

RF equation

fall.equation <- "recent_fall ~ age + full_id + gender + dx"
fall.formula <- as.formula(fall.equation)

fall.model <- randomForest(formula=fall.formula,
             data=pt_data_new,
             ntree = 500,
             mtry = 3,
             nodesize = 0.01*nrow(pt_data_new)
             )

Fall.Prediction <- predict(fall.model,newdata=fall_test_new)

PtID <- fall_test_new$id
output.df <- as.data.frame(PtID)

output.df$Fall.Prediction <- Fall.Prediction

write.csv(output.df,"fall.prediction.csv",row.names = FALSE)

Predicted recent_fall

Fall.Prediction total_percent
<chr>   <chr>
minor fall  44%
no fall 56%

Original dataset recent_fall variables and distribution

recent_fall total_percent
<fct>   <chr>
major fall  19%
minor fall  36%
no fall 46%
0 Answers
Related