Hdf5 file reading speed slowing down suddenly

Viewed 163

I am trying to iterate through a hdf5 file. The file consists of a series of 'timestamps', and for each one there is a variable number of 'players', from which I want to obtain data. The file is quite big though, there are over 17k timestamps, and I observe that the time taken in getting 'players = list(data[timestamp].keys())' suddenly increases a lot after some hundred iterations, from about 0.0005 second to about 0.05.

        with h5py.File(self.hdf5file, "r") as f:
            data = f['data']
            timestamps = list(data.keys())
            for timestamp in timestamps:
                start_time = time.time()
                players = list(data[timestamp].keys())
                end_time = time.time()
                print(end_time - start_time)

I have no idea what might be happening and how to work around it.

1 Answers

I suspect there's more going on than your your simple example captures. It should not show performance degradation when you are simply reading keys to get group and dataset names.

To demonstrate, I wrote an example that creates a file that mimics your schema. After it creates the file, it closes and reopens in read-only mode, then loops to access each timestamp group and player dataset and prints access time. Response time is consistently low for all timesteps. [Example file is 1.4 GB. Increase the size of arr if you want to test a larger file.] FYI, I ran this on a Windows system with 24 GB RAM.

Code below:

ntimes = 10000
nplayers = 10
hdf5file = 'SO_70466356.h5'
start_time = time.time()
with h5py.File(hdf5file, "w") as h5f:
    data = h5f.create_group('data')
    for t_cnt in range(ntimes):
        grp = data.create_group(f'time_{t_cnt:05}')
        for p_cnt in range(nplayers):
            arr = np.random.random(300*3).reshape(300,3)
            grp.create_dataset(f'player_{p_cnt:02}',data=arr)
print(f'HDF5 file creation time: {time.time()-start_time:.4f}')            

print(f'HDF5 file read times')            
with h5py.File(hdf5file, "r") as h5f:
    data = h5f['data']
    start_time = time.time()
    timestamps = list(data.keys())
    for cnt,timestamp in enumerate(timestamps):            
        players = list(data[timestamp].keys())
        if not (cnt+1)%1000 :
            print(f'For {timestamp}: {time.time()-start_time:.4f}')
            start_time = time.time()

Output from above:

HDF5 file creation time: 29.7558
HDF5 file read times
At time_00999: 0.1766
At time_01999: 0.1736
At time_02999: 0.1650
At time_03999: 0.1872
At time_04999: 0.1716
At time_05999: 0.1740
At time_06999: 0.1716
At time_07999: 0.1716
At time_08999: 0.1720
At time_09999: 0.1872
Related