What's the fastest way to return that indices of values of two arrays that are equal to each other?

Viewed 251

Say I have these two numpy arrays:

A = np.array([[1,2,3],[4,5,6],[8,7,3])
B = np.array([[1,2,3],[3,2,1],[8,7,3])

It should return

[0,2]

Since the values at the 0th and 2nd index are equal to each other.

What's the most efficient way of doing this?

I tried something like:

[val for val in range(len(A)) if A[val]==B[val]]

but got the error:

ValueError: The truth value of an array with more than one element is ambiguous. Use a.any() or a.all()
5 Answers

You can do something like that

>>> [a in B for a in A]
[True, False, True]

>>> A[[a in B for a in A]]
array([[1, 2, 3],
       [8, 7, 3]])

>>> np.where((A==B).all(axis=1))
(array([0, 2]),)

Assuming A.shape == B.shape (otherwise just take A=A[:len(B)] and B=B[:len(A)]) consider:

>>> A==B
[[ True  True  True]
 [False False False]
 [ True  True  True]]

>>> (A==B).all(axis=1)
[ True False  True]

>>> np.argwhere((A==B).all(axis=1))
[[0]
 [2]]

The following solution works also for arrays that do not match in their first dimension, i.e., have a different number of rows. It also works if a match occurs multiple times.

import numpy as np
from scipy.spatial import distance


A = np.array([[1, 2 ,3], 
              [4, 5 ,6],
              [8, 7, 3]])

B = np.array([[1, 2, 3],
              [3, 2, 1],
              [1, 2, 3],
              [9, 9, 9]])

res = np.nonzero(distance.cdist(A, B) == 0)
# ==> (array([0, 0]), array([0, 2]))

The result res is a tuple of two array, which represent the match index of the first and the second input array, respectively. So, in this example, the row at the 0th index of the first array matches the row of the 0th index of second array, and the row at the 0th index of the first array matches the row at the second index of the second array.

In [174]: A = np.array([[1,2,3],[4,5,6],[8,7,3]])
     ...: B = np.array([[1,2,3],[3,2,1],[8,7,3]])

Your list comprehension works fine for lists:

In [175]: Al = A.tolist(); Bl = B.tolist()
In [177]: [val for val in range(len(Al)) if Al[val]==Bl[val]]
Out[177]: [0, 2]

For lists == is a simple boolean test - same or not; for arrays it returns a boolean array, which can't be use in an if:

In [178]: Al[0]==Bl[0]
Out[178]: True
In [179]: A[0]==B[0]
Out[179]: array([ True,  True,  True])

With arrays, you need to add a all as suggested by the error:

In [180]: [val for val in range(len(A)) if np.all(A[val]==B[val])]
Out[180]: [0, 2]

The list version will be faster.

But you can also compare the whole arrays, and take row by row all:

In [181]: A==B
Out[181]: 
array([[ True,  True,  True],
       [False, False, False],
       [ True,  True,  True]])
In [182]: np.all(A==B, axis=1)
Out[182]: array([ True, False,  True])
In [183]: np.nonzero(np.all(A==B, axis=1))
Out[183]: (array([0, 2]),)
Related