I am trying to use pandas_udf since my data is in a PySpark dataframe but I would like to use a pandas library. I have a lot of rows so I cannot convert my PySpark dataframe into a Pandas dataframe.
I use textdistance (pip3 install textdistance)
And import it: import textdistance.
test = spark.createDataFrame(
[('dog cat', 'dog cat'),
('cup dad', 'mug'),],
['value1', 'value2']
)
@pandas_udf('float', PandasUDFType.SCALAR)
def textdistance_jaro_winkler(a, b):
return textdistance.jaro_winkler(a, b)
test = test.withColumn('jaro_winkler', textdistance_jaro_winkler(col('value1'), col('value2')))
test.show()
I am getting the following getting error:
ValueError: The truth value of a Series is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
I tried to pass the whole dataframe as an argument in the function and pass string values in the function but I believe it made it worse:
schema = StructType([StructField("value1", StringType(), True)
,StructField("value2", StringType(), True)
,StructField("jaro_winkler", FloatType(), True)
])
@pandas_udf(schema, PandasUDFType.GROUPED_MAP)
def textdistance_jaro_winkler(df):
df['jaro_winkler'] = df.apply(lambda x: textdistance.jaro_winkler(x['value1'], x['value2']))
return df