Efficient way to find closest vector from 10M of samples

Viewed 267

Say I have a database of 10 000 000 of 100-dimensional vectors:

X1 = [x1_1, ..., x1_100]
X2 = [x2_1, ..., x2_100]
...
X1000000 = [x1000000_1, ..., x1000000_100]

And I have input vector Y :

Y = [y1, ..., y100]

What is the most efficient way to find closest vector Xi to Y in sense of euclidean distance?

1 Answers

try this

def find_CD(X,y):
  return ​return​ spatial.distance.euclidean(X,y)

in the main

listt=[]
vic=[x1,x2,.....,X1000000]
for i in range(len(vic)):
 listt.append(find_CD(vic[i],y)

and find the min values index

listt.index(min(listt))
Related