zlib to compress/decompress in a file with python

Viewed 3322

using zlib, i want to be able to compresse numpy arrays and write them in a file and then be able to read them back. I did the following

with open(outputFile, 'wb') as zFile:
    for row in array:
        compressed = zlib.compress(row, compressionLevel)
        zFile.write(compressed)
with open(os.path.join(path, fileName), 'rb') as zFile:
    for line in zFile:
        decompressed = zlib.decompress(line)
        data.append(decompressed)
data = np.array(data)

The writing process works as it fills the file and if i write simpler data with compressionLevel = 0, it is ok. But i can't make the reading process to work. I tried to do to zlib.compress(row.tobytes() + '\n'.encode(), compressionLevel) so that i can have proper lines to be read, but some element in my data seems to be interpreted as \n, so it does not read the real lines.

I also tried to read the file doing zFile.read(bufferSize) in a while loop and break the loop when there is nothing more to read, but each element previously compressed have a varying size (due to varying performance regarding each row) so i can't know the buffersize in advance.

EDIT: regarding answers, it seems that np.savez_compress is better suited but for now, i am stuck with zlib as it could be used elsewhere in the project and i cannot change it by myself for now.

3 Answers

One of the best options to compress numpy arrays is using np.savez_compressed. This will be nicer but will be slower. I don't think your compression code is correct

import numpy as np
import zlib
input_arr = np.arange(100)
dtype = input_arr.dtype
compressed_arr = zlib.compress(input_arr, 2)
decompressed_arr = np.fromstring(zlib.decompress(compressed_arr), dtype)

You can also use blosc which has even better performace

Use the builtin numpy.savez_compressed? From the numpy docs:

>>> test_vector = np.random.rand(4)
>>> np.savez_compressed('/tmp/123', a=test_array, b=test_vector)
>>> loaded = np.load('/tmp/123.npz')
>>> print(np.array_equal(test_array, loaded['a']))
True
>>> print(np.array_equal(test_vector, loaded['b']))
True

so, to sum-up, i cannot use something else than zlib for now but :

  • I can totaly compress/decompress a (unique) numpy array and write/read it to/from a file doing as follow :
row = array[0,:]
with open(outputFile, 'wb') as zFile:
    print(row)
    compressed = zlib.compress(row, compressionLevel)
    zFile.write(compressed)
with open(os.path.join(path, fileName), 'rb') as zFile:
    decompressed = zlib.decompress(zFile.read())
    data = np.frombuffer(decompressed, dtype=np.float))
    print(data)
  • I added np.frombuffer as pointed out by sagarwal (actually np.frombuffer is better than np.fromstring) and print(row) and print(data) gives the same thing. The probleme comes when i add a for loop in the writing process to add several compressed row in the file. thus i have trouble to retrieve each full size compressed row (with a for line in zFile: ...) to decompress them one at a time. Indeed, some element in each compressed row are seen as \n and thus is does not retrieve the real lines (wich gives zlib.error: Error -5 while decompressing data: incomplete or truncated stream)
Related