Scipy.spatial.ckdtree exhaust memory for large data set

Viewed 471

I have two large coordinate set A and B, both defined in the same region.

For each point in A, I want to find out the number of neighbors of type A which the maximum distance to this point is r, and also the number of neighbors of type B. Since my data is large, A contains 262,001 points, and B contains ~340,000 points, I use the cKDTree function from Scipy.spatial.cKDTree to do my job.

I first construct the trees:

import numpy as np
from Scipy.spatial import cKDTree
from Scipy.io import loadmat
from Scipy.stats import spearman's
import math
# load points
CD3_Points = np.array(load mat('./Points_Bivariate/Case 0/CD3_Points.mat)['x])/125
CD4_Points = np.array(load mat('./Points_Bivariate/Case 0/CD4_Points.mat)['x])/125
# construct tree
CD3_Tree = cKDTree(CD3_Points)
CD4_Tree = cKDTree(CD4_Points)

since I am interested in the results as multiple radius distance, I write a for-loop

A_list = np.array([]).reshape(len(CD3_Points), 0)
B_list = np.array([]).reshape(len(CD3_Points), 0)

for r in np.arrange(0.1, 2.8, 0.1):
    A_index = CD3_Tree.query_ball_tree((CD4_Tree, r = r))
    B_index = CD4_Tree.query_ball_tree((CD4_Tree, r = r))
    # get the number
    A_num = np.asarray(list(map(len, A_sub))).reshape(len(CD3_Points),1)
    B_num = np.asarray(list(map(len, B_sub))).reshape(len(CD3_Points),1)
    # append to store data
    A_list = np.append(A_list, A_num, axis = 1) 
    B_list = np.append(B_list, B_num, axis = 1) 

In that case, I assume I could obtain two arrays. For each array, each column represents the search results at one radius for all points.

It went on well initially, however, when r turns to 0.4 I got an error:

Process finished with exit code 137(Interrupted by signal 9: SIGKILL)

I assume it's because the my data sets are two large, so instead of put all points of 'CD4_Points' into the function, I used a for loop to put individual element of CD4_Points to the function, now it looks like:

for its in range(0, len(CD3_Points)):
    Point = CD3_Points[pts,]
    for r in np.arange(0.1, 2.8, 0.1):
        A_num = len(CD3_Tree.query_ball_point(Point, r = r))
        B_num = len(CD4_Tree.query_ball_point(Point, r = r))
        A.append(A, A_num)
        B.append(A, A_num)

    rho, oval = spearmanr(A, B)
    A = []
    B = []

        ...

This seems working but it's really time consuming...is there any solution to make it faster? Besides, I also noticed that there are other types of search methods such as ball tree, and sklearn also has cKDTree and KDTree, which one works better for my case?

0 Answers
Related