I'm looking for a pyspark alternative to the sklearn cosine similarity package

Viewed 36

Here is my data schema:

root
 |-- embeddings: array (nullable = true)
 |    |-- element: float (containsNull = true)
 |-- preprocessed_product_description: string (nullable = true)
 |-- p_id: long (nullable = true)
 |-- id: long (nullable = true)

My ultimate goal is to calculate a cosine similarity between all of the embeddings within a id. Here is the pythonic code:

from sklearn.metrics.pairwise import cosine_similarity
    
cosine_sims = (
       embeddings.groupby("id")
       .apply(lambda x: cosine_similarity(x["embeddings"].tolist()).tolist())
       .reset_index()
    )

Now, I know that won't be possible in pyspark but I would like to create some UDF that will allow me to generate a NxN cosine similarity matrix.

Any help would be much appreciated.

0 Answers
Related