I have a very very large document base, and turn each document into a Set of its ngrams (character based ngrams) and then Use a CountVecotrizer. I want to speed up both the actual distance calculation (by using MinHashes to approixmate Jaccard Distance, but also use a LSH technique for Bucketing. This will give me false negatives and false positives in both the bucketing and the minhash steps of the algorithm, but thats ok. It's the only way I can crunch my data.
My problem is, that sparks MinHash returns an Array(DenseVector, true) where each DenseVector is 1-dim.
LSH then expects a DenseVector. So what I want to do is turn an Array of 1-dim DenseVectors into a n-dim DenseVector. How can I do that with Spark?
I have uncussefully tried
- vectorassembler
- udf
- pandas_udf