Why doesn't CondensedNearestNeighbour() end up with large data?

Viewed 44

I run CondensedNearestNeighbour() undersampling method in jupyter notebook for 1 million rows, according to one variable and target. I think it takes long time. Almost two days are over but, it is still running without result.

I really don't understand, if it doesn't work for huge data, what does it do. I need undersampling to reduce sample number. I don't want to use random sampling. If you have any opinion, i would appreciate. My code sample is below:

X = df1[['var1']].to_numpy()
y=df1['target'].to_numpy()

 
counter = Counter(y)
undersample = CondensedNearestNeighbour(random_state=44, n_neighbors=1)
X1, y1 = undersample1.fit_resample(X, y)
sample_counter = Counter(y1)
0 Answers
Related