All this code works in pandas, but running single threaded is slow.
I have an object (it's a bloom filter) that's slow to create.
I have dask code that looks something like:
def has_match(row, my_filter):
return my_filter.matches(
a=row.a, b =row.b
)
# ....make dask dataframe ddf
ddf['match'] = ddf.apply(has_match, args=(my_filter, ), axis=1, meta=(bool))
ddf.compute()
When I try to run this I get an error that starts:
distributed.protocol.core - CRITICAL - Failed to Serialize
My object was created from a C library, so I'm not surprised that it can't be automagically serialized, but I don't know how to work around this.