I am trying to iterate through a hdf5 file. The file consists of a series of 'timestamps', and for each one there is a variable number of 'players', from which I want to obtain data. The file is quite big though, there are over 17k timestamps, and I observe that the time taken in getting 'players = list(data[timestamp].keys())' suddenly increases a lot after some hundred iterations, from about 0.0005 second to about 0.05.
with h5py.File(self.hdf5file, "r") as f:
data = f['data']
timestamps = list(data.keys())
for timestamp in timestamps:
start_time = time.time()
players = list(data[timestamp].keys())
end_time = time.time()
print(end_time - start_time)
I have no idea what might be happening and how to work around it.