cuda.local.array()
In How is performance affected by using numba.cuda.local.array() compared with numba.cuda.to_device()? a benchmark of the simple quicksort algorithm demonstrates that using to_device to pass preallocated arrays can be ~2x more efficient, but this requires more memory.
The benchmark results for individually sorting 2,000,000 rows each with 100 elements is as follows:
2000000 Elapsed (local: after compilation) = 4.839058876037598 Elapsed (device: after compilation) = 2.2948694229125977 out is sorted Elapsed (NumPy) = 4.541851282119751
Dummy Example using to_device()
If you have a complicated program that has many cuda.local.array() calls, the equivalent to_device version might start to look like this and get quite cumbersome:
def foo2(var1, var2, var3, var4, var5, var6, var7, var8, var9, var10, out):
for i in range(len(var1)):
out[i] = foo(var1, var2, var3, var4, var5, var6, var7, var8, var9, var10, out)
def foo3(var1, var2, var3, var4, var5, var6, var7, var8, var9, var10, out):
idx = cuda.grid(1)
foo(var1, var2, var3, var4, var5, var6, var7, var8, var9, var10, out[idx])
In a real codebase, there might be 3-4 levels of function nesting across tens of functions and hundreds to thousands of lines of code. What are alternatives to these two approaches?