I'm trying to run some code to preprocess my data for machine learning in Caret. One step I'm having a lot of trouble with is KNN imputation. When I run the following block of code:
library(caret)
traindf <- data.frame(matrix( rnorm(7*7,mean=0,sd=1), nrow=7, ncol=7))
testdf <- data.frame(matrix( rnorm(7*7,mean=0,sd=1), nrow=7, ncol=7))
for(i in 1:7){
traindf[i,i] <- NA #generates NA's in every row
}
impute_model <- preProcess(traindf, method = c('knnImpute')) #this line is problematic
imputed_train <- predict(impute_model, traindf)
imputed_test <- predict(impute_model, testdf)
I get an error:
Error in RANN::nn2(old[, non_missing_cols, drop = FALSE], new[, non_missing_cols, :
Cannot find more nearest neighbours than there are points
From some research, I believe this is due to the fact that the kNN imputation implementation Caret uses discards rows with any NA's. In my dataset, NA's are scattered throughout such that this would result in all rows being discarded for imputation purposes. Instead I would like to keep these partially NA rows and still use them for imputation.
I know of one package that does this:https://www.rdocumentation.org/packages/impute/versions/1.46.0/topics/impute.knn. However, this one doesn't override predict, so I can't use it easily to impute the test set as well like in the above example.
Does anyone have suggestions on how I can get this partial-NA KNN imputation working with Caret?