I am looking to compute the 2D KDE of some large dataset on Python.
Whereas fitting the KDE performs well both in R (using ks package) and Python (whatever the package), evaluating the density for some new random large dataset is instantaneous in R, but takes a long in Python.
Is there anything I'm missing ? How can I speed up things in Python without parallelizing calculations ? Is there a significant difference in the way R and Python algorithms are coded ?
You will find below a reproducible example in R.
library(ks)
library(MASS)
data <- mvrnorm(1e5, mu = c(0, 0), Sigma = matrix(c(10, 3, 3, 2), nrow = 2))
kn <- kde(data, H = diag(2))
new_data <- mvrnorm(1e5, mu = c(1, 5), Sigma = matrix(c(20, 6, 6, 4), nrow = 2))
a <- Sys.time()
kn_pred <- predict(kn, x = new_data)
b <- Sys.time()
b-a #Time difference of 0.005482912 secs
And now the same example in Python.
import time
import numpy as np
from scipy.stats import gaussian_kde
data = np.random.multivariate_normal(mean = [0,0], cov = [[10, 3], [3, 2]], size = 100000)
kn = gaussian_kde([data[:,0], data[:,1]], bw_method = 1)
new_data = np.random.multivariate_normal(mean = [1,5], cov = [[20, 6], [6, 4]], size = 100000)
a = time.time()
kn_pred = kn([new_data[:,0], new_data[:,1]])
b = time.time()
print(b-a) # 131.60702800750732
Thanks in advance for your help.