I have a list of Python objects that I want to cluster into an unknown number of groups. The objects can not simply be compared by any distance function proposed by scikit-learn, but rather by a custom defined one. I'm using DBSCAN from the scikit-learn library, which when run on my data raises a TypeError.
Here's what the faulty code looks like. The objects I want to cluster are "Patch" objects, obtained from scanning a 3d mesh :
from sklearn.cluster import DBSCAN
def getPatchesSimilarity(patch1, patch2):
... #Logic to calculate distance between patches
return dist
#Reading the data (a mesh object) and extracting its patches
mesh = readMeshFromFile("foo.obj")
patchesList = extractPatchesFromMesh(mesh)
clustering = DBSCAN(metric = getPatchesSimilarity).fit(np.array([[patch] for patch in meshPatches]))
When run, this code produces the following error :
TypeError: float() argument must be a string or a number, not 'Patch'
Which seems to mean that the DBSCAN algorithm as proposed by scikit-learn doesn't work with values that aren't vectors or strings ?
I have tried also to use only the indices of the patches, so that the data passed was numerical, but it also didn't work. The last solution that would work now would be to use a distance matrix, but the number of objects is really large, and my computer wouldn't be able to store such a matrix.