I update this question based on the comments as it may be useful to other people
I had a bottleneck in an algorithm I was developing:
eval(parse(text = XXX))
The outcome I had were millions of text strings with some condition I needed to evaluate.
My question was that eval(parse(text = XXX)) is 100 times slower than if the same condition is called in the console.
As an example:
par_cond <- parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766")
microbenchmark(console = 5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766,
`eval(parse` = eval(parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766")),
keep_source = eval(parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766" , keep.source = F)),
par_cond = eval(par_cond))
Unit: nanoseconds
expr min lq mean median uq max neval cld
console 400 501 538.00 501 601.0 1201 100 a
eval(parse 51101 51851 54309.04 52101 52551.5 122200 100 c
keep_source 18901 19301 19674.13 19601 19901.0 22500 100 b
par_cond 1400 1501 1743.86 1701 1801.0 7401 100 a
QUESTION: Can you suggest any way to make this faster?
MY ANSWER:
instead of
eval(parse(text = XXX))
do
# outside the loop
XXX_parsed <- parse(text = XXX)
# inside the loop
eval(XXX_parsed)
CONTEXT:
The algorithm I'm developing is a classification one. It consider the position in space and evaluate if a condition is met based on certain conditions and variables. It is very flexible in the sense that an user can input what variables and condition he wants.
The conditions are inputted as text string and the numbers you see above are the results of replacements of text (referring to variables stored in a hash tables) with numbers.
Most of these conditions cannot be evaluated at once as the algorithm starts on a few anchor points and consider only objects adjacent that meet conditions. It keeps on with the evaluation as long as there are new objects adjacent that meet conditions.
Examples of conditions that can be evaluated:
condA <- "light > 1000 & between(oxygen, 0.5, 0.8) & elevation < elevation[]"
condB <- "temperature <= 2 & (bias == F | Si < 0.2) & surface > surface[]"
Note that the [] (e.g. elevation[]) means that the elevation of the object considered (e.g. elevation) is tested against one or more objects next to it (e.g. elevation[]).
UPDATED APPROACH:
After posting this question I was suggested to parse whatever condition outside loops and run only the eval() command at each iteration.
so for instance imagine the inputs:
cond <- "light > 1000 & between(oxygen, 0.5, 0.8) & elevation < elevation[]"
var <- data.table(light = 995:1005, oxygen = (runif(11)), elevation = sample(500:550, 11))
The way I implemented it now is to create two lists with all the values needed to evaluate my expression:
v_ab <- names(var)[str_detect(cond, paste0(names(var), "(?!\\[|\\{)"))]
v_fn <- names(var)[str_detect(cond, paste0( names(var),"\\[\\]" ))]
for(v in v_ab){cond <- str_replace_all(cond, paste0(v, "(?!\\[)"), paste0("l_ab$", v))}
for(v in v_fn){cond <- str_replace_all(cond, paste0(v, "\\[\\]" ), paste0("l_fn$", v))}
cond_parsed <- parse(text = cond)
> cond_parsed
expression(l_ab$light > 1000 & between(l_ab$oxygen, 0.5, 0.8) & l_ab$elevation < l_fn$elevation)
The values of l_ab$light, l_ab$oxygen, l_ab$elevation and l_fn$elevation are updated at each iteration and the conditions are evaluated calling
eval(cond_parsed)