Reading Empty CSV Pyspark

Viewed 808

I have a process to read csv files and do some processing in pyspark. At times I might get a zero byte empty file. In such cases when I use the below code

df = spark.read.csv('/path/empty.txt', header = False)

It is failing with error:

py4j.protocol.Py4JJavaError: An error occurred while calling o139.csv. : java.lang.UnsupportedOperationException: empty collection

Since its empty file I tried to read as a json it worked fine

df = spark.read.json('/path/empty.txt')

When I add header to the empt csv manually to the code reads fine.

df = spark.read.csv('/path/empty.txt', header = True)

In few places I read to use databricks csv but I don't have the data bricks csv package options to use as those jars are not available in my environment.

0 Answers
Related