I use this apply function on my code:
def entities_extraction(text):
doc = nlp(text)
entities= [ent.text for sentence in doc.sentences for ent in sentence.entities if ent.type in {"PERSON", "ORG", "GPE", "NORP", "FAC", "LOC", "PRODUCT", "EVENT", "WORK_OF_ART", "LAW", "LANGUAGE", "MISC"}]
return entities
df["entities"] = df['text'].progress_apply(lambda x: entities_extraction(x))
The problem is that, at the moment, it is too slow (it take nearly 12 hours)
So I tried to modify it to use colab gpu:
@cuda.jit
def entities_extraction(text):
doc = nlp(text)
entities= [ent.text for sentence in doc.sentences for ent in sentence.entities if ent.type in {"PERSON", "ORG", "GPE", "NORP", "FAC", "LOC", "PRODUCT", "EVENT", "WORK_OF_ART", "LAW", "LANGUAGE", "MISC"}]
return entities
But I get this error:
ValueError:
Kernel launch configuration was not specified. Use the syntax:
kernel_function[blockspergrid, threadsperblock](arg0, arg1, ..., argn)
See https://numba.pydata.org/numba-doc/latest/cuda/kernels.html#kernel-invocation for help.
Do you know how to solve it, or you have a better implementation for gpu on apply functions?
Sorry if I made a lot of mistakes, I'm new to the argument of gpu for faster code. Thank you for your help!