How to convert distance based algorithm to points-based DBscan algorithm in Clustering.jl for outlier detection

Viewed 51

I am trying to detect outliers for a large dataset using DBscan. However due to its huge size, it is giving out of memory error. In documentation of DBSCAN ( https://juliastats.org/Clustering.jl/stable/dbscan.html ) it is telling that points based algo takes much less memory and is efficient.

However I am not understanding properly how to convert it to points based format. I want the same output from points based approach.

Code for distance based method

using Clustering
using Statistics
using Distances

d = randn( 1000 );

c = vcat( d, [ 9999, 8888 ] );

dist = pairwise(Euclidean(), c'; dims=2);

dbout = dbscan( dist, 1, 5 );

isoutlier = 1 .- dbout.assignments
0 Answers
Related