Here is my data schema:
root
|-- embeddings: array (nullable = true)
| |-- element: float (containsNull = true)
|-- preprocessed_product_description: string (nullable = true)
|-- p_id: long (nullable = true)
|-- id: long (nullable = true)
My ultimate goal is to calculate a cosine similarity between all of the embeddings within a id. Here is the pythonic code:
from sklearn.metrics.pairwise import cosine_similarity
cosine_sims = (
embeddings.groupby("id")
.apply(lambda x: cosine_similarity(x["embeddings"].tolist()).tolist())
.reset_index()
)
Now, I know that won't be possible in pyspark but I would like to create some UDF that will allow me to generate a NxN cosine similarity matrix.
Any help would be much appreciated.