I am trying to reduce the memory consumption of a piece of R code I have been working on. I am using the peakRAM() function to measure the maximum RAM used. It is a long code and there is a simple sapply() function at the end of it. I figured out that it is the sapply() part which is consuming the maximum memory. So I have written a small function fun1() imitating the objects and the sapply() function from that part of my code, which is as follows :
library(peakRAM)
fun1 <- function() {
tm <- matrix(1, nrow = 300, ncol = 10) #in the original code, the entries are different and nonzero
print(object.size(tm))
r <- sapply(1:20000, function(i) {
colSums(tm[1:200,]) #in the original code, I am subsetting a 200 length vector which varies with i, stored in a list of length 20000
})
print(object.size(r))
r
}
peakRAM(fun1())
If you run this in R, you get a peakRAM() consumption of around 330Mb. But you can see that the two objects tm and r are both of very small size (2Kb and 1.6Mb respectively) and if you look at the peakRAM() for computing a single colSums(tm[1:200,]), it is very small, like 0.1Mb. So it feels like, during sapply(), R is probably not getting rid of the memory while looping over 1:20000. Otherwise, since a single colSums(tm[1:200,]) takes very small memory, and all the objects associated are of small memory, the sapply() should have taken small memory.
In this regard, I already know that R has a gc() function which gets rid of unnecessary memory when needed and probably R is not clearing memory during sapply() which is resulting into this high memory consumption. If that is true, I would like to know if there is a way to get rid of this and complete the job without requiring this much extra memory? Note that, I do not wish to compromise on the run-time for doing that.