I am working with PySpark on Google Colab, and after sometime, I am keep getting an error when I perform tasks like df.show(), train and test split using splits = df.randomSplit([0.8, 0.2]) in pyspark, and many more. This problem also happens when I shuffle the dataframe. Sometimes it gives an error saying memory overflow, and sometimes it displays this error:
ERROR:root:Exception while sending command.
Traceback (most recent call last):
File "/usr/local/lib/python3.7/dist-packages/IPython/core/interactiveshell.py", line 3326, in run_code
exec(code_obj, self.user_global_ns, self.user_ns)
File "<ipython-input-32-4d37eba09562>", line 1, in <module>
df4.show()
File "/content/spark-3.3.0-bin-hadoop3/python/pyspark/sql/dataframe.py", line 606, in show
print(self._jdf.showString(n, 20, vertical))
File "/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/java_gateway.py", line 1322, in __call__
answer, self.gateway_client, self.target_id, self.name)
File "/content/spark-3.3.0-bin-hadoop3/python/pyspark/sql/utils.py", line 190, in deco
return f(*a, **kw)
File "/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/protocol.py", line 328, in get_return_value
format(target_id, ".", name), value)
py4j.protocol.Py4JJavaError: <unprintable Py4JJavaError object>
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/usr/local/lib/python3.7/dist-packages/IPython/core/interactiveshell.py", line 2040, in showtraceback
stb = value._render_traceback_()
AttributeError: 'Py4JJavaError' object has no attribute '_render_traceback_'
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/clientserver.py", line 516, in send_command
raise Py4JNetworkError("Answer from Java side is empty")
py4j.protocol.Py4JNetworkError: Answer from Java side is empty
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/java_gateway.py", line 1038, in send_command
response = connection.send_command(command)
File "/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/clientserver.py", line 540, in send_command
"Error while sending or receiving", e, proto.ERROR_ON_RECEIVE)
py4j.protocol.Py4JNetworkError: Error while sending or receiving
---------------------------------------------------------------------------
Py4JJavaError Traceback (most recent call last)
/usr/local/lib/python3.7/dist-packages/IPython/core/interactiveshell.py in run_code(self, code_obj, result, async_)
3325 else:
-> 3326 exec(code_obj, self.user_global_ns, self.user_ns)
3327 finally:
13 frames
<class 'str'>: (<class 'ConnectionRefusedError'>, ConnectionRefusedError(111, 'Connection refused'))
During handling of the above exception, another exception occurred:
ConnectionRefusedError Traceback (most recent call last)
/content/spark-3.3.0-bin-hadoop3/python/lib/py4j-0.10.9.5-src.zip/py4j/clientserver.py in connect_to_java_server(self)
436 self.socket = self.ssl_context.wrap_socket(
437 self.socket, server_hostname=self.java_address)
--> 438 self.socket.connect((self.java_address, self.java_port))
439 self.stream = self.socket.makefile("rb")
440 self.is_connected = True
ConnectionRefusedError: [Errno 111] Connection refused
The operation df.show() works properly at the begining, but after performing some operations, it keeps giving the errors.
Some people were saying to increase the drive size. Like in this case:
py4j.protocol.Py4JNetworkError: Answer from Java side is empty while trying to execute df.show(5)
And therefore, I increased the driver size but that too did not work. I did the following configuration for that:
from pyspark import SparkContext
SparkContext.setSystemProperty('spark.executor.memory', '16g')
SparkContext.setSystemProperty("spark.driver.memory", "16g")
sc = SparkContext("local", "Classification").getOrCreate()
from pyspark.sql import SparkSession
spark = (SparkSession.builder.appName("Classification").getOrCreate())
I tried running this on local machine, and it is giving the same errors.
I tried running this code:
from pyspark.sql import SparkSession
spark = SparkSession.builder.appName("Classification").getOrCreate()
spark.conf.set("spark.executor.memory.pyspark", "16g")
But this too gives the same error.
Update 00: I tried this code in colab,
from pyspark.sql import SparkSession
spark = SparkSession.builder\
.appName("test")\
.config("spark.driver.memory", "22g")\
.config("spark.executor.memory", "22g")\
.getOrCreate()
sc = spark.sparkContext
from pyspark.sql import SQLContext
sqlContext = SQLContext(sc)
But it is taking little more time then before and then it shows the same error. I increase the memory from 22g to 100g and then it is taking too much time to execute and I have to pause it.
Then in local machine I tried this command,
spark-shell --driver-memory 100g
And it too was taking too much time and I have to pause the execution
Could someone please help me solve this issue?