I am trying to detect outliers for a large dataset using DBscan. However due to its huge size, it is giving out of memory error. In documentation of DBSCAN ( https://juliastats.org/Clustering.jl/stable/dbscan.html ) it is telling that points based algo takes much less memory and is efficient.
However I am not understanding properly how to convert it to points based format. I want the same output from points based approach.
Code for distance based method
using Clustering
using Statistics
using Distances
d = randn( 1000 );
c = vcat( d, [ 9999, 8888 ] );
dist = pairwise(Euclidean(), c'; dims=2);
dbout = dbscan( dist, 1, 5 );
isoutlier = 1 .- dbout.assignments