Inside numba .jit(nopython=True) function i am calculating thousands of numpy arrays (1-D, integer data type) and append them to the list. The problem is that some of the arrays appears equal, but i dont need duplicates. So i need an efficient way to check if new array already exists in the list or not.
In python it can be done like this:
import numpy as np
import numba as nb
# @nb.jit(nopython=True)
def foo(n):
uniques = []
uniques_set = set()
for _ in range(n):
arr = np.random.randint(0, 2, 2)
arr_hashable = make_hashable(arr)
if not arr_hashable in uniques_set:
uniques_set.add(arr_hashable)
uniques.append(arr)
return uniques
Ive tried two ways to solve this:
Converting array to tuple and put the tuple inside of a set.
def make_hashable(arr): return tuple(arr)but unfortunately direct tuple construction doesnt work this way in nopython mode. Ive tried also this way:
def make_hashable(arr): res = () for n in arr: res += (n,) return resand other similar workarounds i could think of, but all of them failed in nopython mode with TypeError.
Convert array to string and also put it to the set.
def make_hashable(arr): return arr.tostring()also tried all possible ways to convert array to string but seems like numba doesnt support string conversion for now
Maybe there are different approaches to check (efficiently) if array is already exists in a list? My numba version is 0.44. Thanks a lot.