I want to run sparknlp in python, I am using apache-spark 3.2.1, spark-nlp==3.4.1 pyspark==3.1.2. I am following this guide. I am able to get the spark session using this code :
sc = pyspark.SparkContext().getOrCreate()
import sparknlp
sparknlp.start()
Whenever I try to download any pre-trained model using code :
pipeline = PretrainedPipeline('explain_document_dl', lang='en')
I get a few errors, I resolved some errors one by one by adding the jar for that error in the apache-spark jars. for eg : One of the errors was :
java.lang.NoClassDefFoundError: org/tensorflow/ndarray/NdArray
which I resolved by adding the NdArray Jar
Like this I added 6-7 jars depending on the error.
The error that I am stuck at is this :
Py4JJavaError: An error occurred while calling z:com.johnsnowlabs.nlp.pretrained.PythonResourceDownloader.downloadPipeline.
: java.lang.VerifyError: Bad return type
Exception Details:
Location:
com/johnsnowlabs/ml/tensorflow/TensorResources.createTensor(Ljava/lang/Object;)Lorg/tensorflow/Tensor; @370: areturn
Reason:
Type 'java/lang/Object' (current frame, stack[0]) is not assignable to 'org/tensorflow/Tensor' (from method signature)
P.S. I am using java 8