I have a large data with more than 1 billion observations, and I need to perform some string operations which is slow.
My code is as simple as this:
DT[, var := some_function(var2)]
If I'm not mistaken, data.table uses multithread when it is called with by, and I'm trying to parallelize this operation utilizing this. To do so, I can make an interim grouper variable, such as
DT[, grouper := .I %/% 100]
and do
DT[, var := some_function(var2), by = grouper]
I tried some benchmarking with a small sample of data, but surprisingly I did not see a performance improvement. So my questions are:
- Does
data.tableuse multithreading when it's used withby? - If so, is there a condition that multithreading is enabled / disabled?
- Is there a way that user can "enforce"
data.tableto use multithreading here?
FYI, I see that multithreading enabled with half of my cores when I import data.table, so I guess there's no openMP issue here.