I am working with a database of 500 columns and 20,000 rows and I want to change the NA data by the statistical mode, so I avoid eliminating those values, simply change it by the mode of the specific column, so I am given an example base to show the code that I am running
library(tidyverse)
temp <- c(20.37, 18.56, NA, 21.96, 29.53, 28.16,
36.38, 36.62, 40.03, 27.59, 22.15, 19.85)
humedad <- c(88, 86, 81, 79, 80, 78,
71, NA, 78, 82, 85, 83)
precipitaciones <- c(72, 33.9, 37.5, 36.6, 31.0, 16.6,
1.2, 6.8, 36.8, 30.8, 38.5, 22.7)
precipitaciones2 <- c(72,NA, 6.8, 36.6, 31.0, 16.6,
1.2, 6.8, 36.8, 6.8, 38.5, 22.7)
precipitaciones3 <- c(72,NA, 37.5, 36, 2, 16.6,
1.2, 8, 0.8, NA, 38.5, 8)
mes <- c("enero", "febrero", "marzo", "abril", "mayo", "junio",
"julio", "agosto", "septiembre", "octubre", "noviembre", "diciembre")
datos <- data.frame(mes = mes, temperatura = temp, humedad = humedad,
precipitaciones = precipitaciones,
precipitaciones2 = precipitaciones2,
precipitaciones3 = precipitaciones3)
I want to replace the NA data with the statistical mode for a much larger database, so what is required is to program it for any other database, I have the following code:
#mode
mode=getmoda<-function(v){
uniqv<-unique(v)
uniqv[which.max(tabulate(match(v,uniqv)))]
}
reemplazar<-function(y){
i=2
lista_vacia1 <- list()
lista_vacia2<-list()
a<-""
while(i<=5){
lista_vacia1<-y[,i] #select the column to filter
lista_vacia2<-lista_vacia1[!is.na(lista_vacia1)] #remove the NA data
a<-mode(lista_vacia2) #get the mode of the column
y<-y %>% mutate_at(i,~replace_na(.,a))
a<- ""
lista_vacia1 <- list()
lista_vacia2<-list()
}
}
What happens is that when I run the program it makes an infinite loop, it never goes beyond loading and it does not show any message. I would like you to help me to know why this happens or if it is possible to change the code.