I want to calculate the Median absolute deviation (mscore) by column ignoring the first column for each dataframe in a list of dataframes. Then add the result as a new row into the dataframe with the row name mscore.
Previously I would do the calculation on each dataframe one at a time but now its streamline the process.
A small extract of my list of dataframes below. The full list of dfs has over 30 dataframes
list(Al2O3 = structure(list(Determination_No = 1:6, `2` = c(2.01,
2.02, 2.03, 2.01, 2.02, 2), `3` = c(2.01, 2.01, 2, 2.02, 2.02,
2.03), `4` = c(2, 2.03, 1.99, 2.01, 2.01, 2.01), `5` = c(2.02,
2.02, 2.05, 2.03, 2.02, 2.03), `7` = c(1.88, 1.9, 1.89, 1.88,
1.88, 1.87), `8` = c(2.053, 2.044, 2.041, 2.038, 2.008, 2.02),
`10` = c(2.002830415, 2.021725042, 2.021725042, 1.983935789,
2.002830415, 2.021725042), `12` = c(2.09, 2.05, 1.96, 2.09,
2.06, 2.02)), class = "data.frame", row.names = c(NA, -6L
)), As = structure(list(Determination_No = 1:6, `2` = c(0.052,
0.027, 0.011, 0.011, 0.012, 0.012), `3` = c(0.012, 0.012, 0.013,
0.012, 0.013, 0.013), `4` = c(0.012, 0.012, 0.013, 0.012, 0.012,
0.012), `5` = c(0.013, 0.013, 0.013, 0.013, 0.013, 0.013), `7` = c(0.011,
0.011, 0.011, 0.012, 0.011, 0.011), `8` = c(0.011, 0.01, 0.011,
0.011, 0.011, 0.011), `10` = c(0.01, 0.01, 0.01, 0.01, 0.01,
0.01), `12` = c(NA_real_, NA_real_, NA_real_, NA_real_, NA_real_,
NA_real_)), class = "data.frame", row.names = c(NA, -6L)), Fe = structure(list(
Determination_No = 1:6, `2` = c(55.94, 55.7, 56.59, 56.5,
55.98, 55.93), `3` = c(56.83, 56.54, 56.18, 56.5, 56.51,
56.34), `4` = c(56.39, 56.43, 56.53, 56.31, 56.47, 56.35),
`5` = c(56.32, 56.29, 56.31, 56.32, 56.39, 56.32), `7` = c(56.48,
56.4, 56.54, 56.43, 56.73, 56.62), `8` = c(56.382, 56.258,
56.442, 56.258, 56.532, 56.264), `10` = c(56.3, 56.5, 56.2,
56.5, 56.7, 56.5), `12` = c(56.11, 56.46, 56.1, 56.35, 56.36,
56.37)), class = "data.frame", row.names = c(NA, -6L)))
Previously I would do the following
#create a modified scores function to accept NAs
scores_na <- function(x, ...) {
not_na <- !is.na(x)
scores <- rep(NA, length(x))
scores[not_na] <- outliers::scores(na.omit(x), ...)
scores
}
MscoreMax <- 3.0 # the the threshold to remove values deemed to be an outlier
colmedians <- median, df[-1], na.rm = T)
MScore <- as.vector(round(abs(scores_na(colmedians, "mad")), digits = 2)) #Mscore to 2 decimals
places
MscoreIndex <- which(MScore > MscoreMax) #get the index of each value exceeding the threshold
df[-1][Fe.MscoreIndex] <- NA # change outliers to NA so they are excluded from further calculations
I have tried the line below to calculate the median
the colmedians function is for a matrix so I have used mapply to apply across the columns
df <- lapply(df, function(x) rbind(x[,-1],
mapply(median(x[,-1],na.rm = TRUE))))
however I get the follow error
Error in median.default(x[, -1], na.rm = TRUE) : need numeric data
when I query the dataframes i the values are stored as double so a bit stuck.