Does anyone know any simple algorithm in Python / PySpark to detect outliers in K-means clustering and to create a list or data frame of those outliers? I'm not sure how to obtain the centroids. I am using the following code:
n_clusters = 10
kmeans = KMeans(k = n_clusters, seed = 0)
model = kmeans.fit(Data.select("features"))