Loop Aggregate with Weighted Mean in R

Viewed 94

Apologies in advance for wording, English is not my native language and this is my first post. I have been able to aggregate my data to this point, but am having issues condensing it further. I am trying to get the weighted average depth by biomass of several species.

My data currently has columns (station, time, layer, depth, biomass_X, biomass_Y, biomass_Z, ...) and I want to condense it to (station, time, weighted_depth_X, weighted_depth_Y, weighted_depth_Z, ...).

I got this code to work, but is there a way to loop it so it can complete all my columns?

    library(plyr)
    newData<-ddply(data, ~station+time, summarize, weighted.mean(data[,6], w=depth))

Data Table, formatting not look right when post

1 Answers

There is certainly a nicer way but this should work:

# data: dataframe containing columns to be averaged
# weights: vector containing the corresponding weights
weighted_mean_all_cols<- function(data,weights){
  res<-do.call(cbind,llply(colnames(data), function(col) {weighted.mean(data[,col], w=weights)}))
  colnames(res) <- colnames(data)
  res
}

# collect the names of the target columns to average
targetCols <-  grep("^biomass",colnames(data))
# apply weighted average by group, for every target column
newData <- ddply(data, c('station','time'), function(groupDF) { 
  print(groupDF[targetCols])
  weighted_mean_all_cols(groupDF[,targetCols],groupDF$depth)
})
Related