Remember: eval(XXX_parsed) is much faster than eval(parse(text = XXX))

Viewed 292

I update this question based on the comments as it may be useful to other people

I had a bottleneck in an algorithm I was developing:

eval(parse(text = XXX))

The outcome I had were millions of text strings with some condition I needed to evaluate.

My question was that eval(parse(text = XXX)) is 100 times slower than if the same condition is called in the console.

As an example:

par_cond <- parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766")
microbenchmark(console = 5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766, 
               `eval(parse` = eval(parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766")),
               keep_source = eval(parse( text = "5.1019344329834 <= 5 & 477 > 0 & -92.8552017211914 < -86.5204010009766" , keep.source = F)),
               par_cond = eval(par_cond))


Unit: nanoseconds
        expr   min    lq     mean median      uq    max neval cld
     console   400   501   538.00    501   601.0   1201   100 a  
  eval(parse 51101 51851 54309.04  52101 52551.5 122200   100   c
 keep_source 18901 19301 19674.13  19601 19901.0  22500   100  b 
    par_cond  1400  1501  1743.86   1701  1801.0   7401   100 a 

QUESTION: Can you suggest any way to make this faster?

MY ANSWER:

instead of

eval(parse(text = XXX))

do

# outside the loop
XXX_parsed <- parse(text = XXX)

# inside the loop
eval(XXX_parsed)

CONTEXT:

The algorithm I'm developing is a classification one. It consider the position in space and evaluate if a condition is met based on certain conditions and variables. It is very flexible in the sense that an user can input what variables and condition he wants.

The conditions are inputted as text string and the numbers you see above are the results of replacements of text (referring to variables stored in a hash tables) with numbers.

Most of these conditions cannot be evaluated at once as the algorithm starts on a few anchor points and consider only objects adjacent that meet conditions. It keeps on with the evaluation as long as there are new objects adjacent that meet conditions.

Examples of conditions that can be evaluated:

condA <- "light > 1000 & between(oxygen, 0.5, 0.8) & elevation < elevation[]"
condB <- "temperature <= 2 & (bias == F | Si < 0.2) & surface > surface[]"

Note that the [] (e.g. elevation[]) means that the elevation of the object considered (e.g. elevation) is tested against one or more objects next to it (e.g. elevation[]).

UPDATED APPROACH:

After posting this question I was suggested to parse whatever condition outside loops and run only the eval() command at each iteration.

so for instance imagine the inputs:

cond <- "light > 1000 & between(oxygen, 0.5, 0.8) & elevation < elevation[]"
var <- data.table(light = 995:1005, oxygen = (runif(11)), elevation = sample(500:550, 11))

The way I implemented it now is to create two lists with all the values needed to evaluate my expression:

v_ab <- names(var)[str_detect(cond, paste0(names(var), "(?!\\[|\\{)"))]
v_fn <- names(var)[str_detect(cond, paste0( names(var),"\\[\\]" ))]

for(v in v_ab){cond <- str_replace_all(cond, paste0(v, "(?!\\[)"), paste0("l_ab$", v))}
for(v in v_fn){cond <- str_replace_all(cond, paste0(v, "\\[\\]" ), paste0("l_fn$", v))}

cond_parsed <- parse(text = cond)

> cond_parsed
expression(l_ab$light > 1000 & between(l_ab$oxygen, 0.5, 0.8) & l_ab$elevation < l_fn$elevation)

The values of l_ab$light, l_ab$oxygen, l_ab$elevation and l_fn$elevation are updated at each iteration and the conditions are evaluated calling

eval(cond_parsed)
0 Answers
Related