Compress CSV to ZIP in dbfs (databricks file storage)

Viewed 139

I'm trying to compress a csv, located in an azure datalake, to zip. The operation is done with python code in databricks, where I created a mount point to relate directly dbfs with the datalake.

This is my code:

import os
import zipfile 

csv_path= '/dbfs/mnt/<path>.csv'
zip_path= '/dbfs/mnt/<path>.zip' 

with zipfile.ZipFile(zip_path, 'w') as zip:
    zip.write(csv_path)  # zipping the file

But I'm getting this error:

OSError: [Errno 95] Operation not supported

Is there any way of doing it?

Thank you in advance.

2 Answers

No, this is not possible to do like you did. The main reason is that local DBFS API has limitations - it doesn't support random writes that is required when you're creating a zip file.

The workaround would be following - output zip file to the local disk of the driver node, and then use dbutils.fs.mv to move file to DBFS, something like this:

import os
import zipfile 

csv_path= '/dbfs/mnt/<path>.csv'
zip_path= '/dbfs/mnt/<path>.zip' 
local_path = '/tmp/my_file.zip'

with zipfile.ZipFile(local_path, 'w') as zip:
    zip.write(csv_path)  # zipping the file
dbutils.fs.mv(f"file:{local_path}", zip_path)

I got same error below while reproducing this.

enter image description here

But I am able to compress csv file into zip by converting to dataframe and then to zip like below.

df=spark.read.csv("dbfs:/mnt/ok/csv1.csv")
df.coalesce(1).write.option("compression","gzip").csv("/dbfs/mnt/ok/myzip2.zip")

enter image description here

Please don't confuse with the path of csv above, here by mistake I have used another csv from ADLS.

You can see the zip file in dbfs below.
enter image description here

But coalesce gives the file name in the zip as part names. To rename it use dbuits.fs.mv(old_path,new_path)

  • First get the csv file path using ls
    enter image description here
  • Then use this path to rename to a new path like below.
old_name = r"/dbfs/mnt/ok/myzip2.zip/part-00000-tid-1285084120372550072-c8b0b7bd-b3b4-4432-8575-4e33e5328ec9-6-1-c000.csv.gz"
new_name=r"/dbfs/mnt/ok/myzip2.zip/mycsv.csv.gz"
dbutils.fs.mv(old_name, new_name)



The above code is referred from this thread by Alex Ott.

After renaming:
enter image description here

Related