hdf5 'Unable to open object (bad object header version number)' error handling

Viewed 1682

I'm reading data from a h5 file that apparently is faulty. It consist of several hundreds of cycles(each with 30members) and one of the cycles(#100) is empty (the group exists but it contains no group members).

When iterating through the cycles, the moment that cycle is reached i get a 'Unable to open object (bad object header version number)' How can i prevent that my script stops when this cycle is reached?

I tried checking in advance whether each group actually has members in order to exclude empty cycles from the iteration but i get a runtimeError already trying to do so:

g = h5py.File('file.h5', 'r')
g['cycle_100']
*** RuntimeError: Can't determine # of objects (bad symbol table node signature)

searching for a solution i found that only exception can be handled not errors - is there nothing i can do but manually exclude this cycle after the error has happened and run the script again? This would be easy but whenever i receive a faulty file i would have to do it again.

Any pointers what i should search for would be appreciated I'm a beginner.

1 Answers

To be clear, accessing a group ('cycles') with no datasets ('members') will not cause an error. The error message (bad symbol table node signature) sounds like something is corrupted in the file. These kinds of errors are hard to reproduce, so this is difficult to diagnose without file. That said, here are some ideas to try.

First, what happens when you run this Python/h5py code:

with h5py.File('file.h5', 'r') as g:
    print(len(g['cycle_1'].keys())) # get number of datasets for a "good" group
    print(len(g['cycle_100'].keys())) # get number of datasets for the "bad" group

I expected you will get a similar error on g['cycle_100'], but it's worth a try.

You can inspect your file with the HDF5 utilities h5ls and h5dump. I think they are delivered with h5py. I found them in my Python's Library\bin folder. (If not, you will need to install HDF5 to get them.) You run these commands on the command line (not as Python code).

  • Use the h5ls command to get a list of all groups and dataset names & shapes, like this: h5ls -r file.h5
  • Use the h5dump command to list header info for all groups and datasets, like this: h5dump -H file.h5.
  • Both commands have other options you may want to investigate. Enter the command name without any parameters to get the 'help' output.

Example output from both utilities for a simple file is shown below. Group group1 has 5 datasets and group2 has no datasets.

E:[.\StackOverflow]->h5ls -r file.h5
/                        Group
/group1                  Group
/group1/dset_01          Dataset {1, 100, 20, 20}
/group1/dset_02          Dataset {1, 100, 20, 20}
/group1/dset_03          Dataset {1, 100, 20, 20}
/group1/dset_04          Dataset {1, 100, 20, 20}
/group1/dset_05          Dataset {1, 100, 20, 20}
/group2                  Group

E:[.\StackOverflow]->h5dump -H file.h5
HDF5 "file.h5" {
GROUP "/" {
   GROUP "group1" {
      DATASET "dset_01" {
         DATATYPE  H5T_IEEE_F64LE
         DATASPACE  SIMPLE { ( 1, 100, 20, 20 ) / ( 1, 100, 20, 20 ) }
      }
      DATASET "dset_02" {
         DATATYPE  H5T_IEEE_F64LE
         DATASPACE  SIMPLE { ( 1, 100, 20, 20 ) / ( 1, 100, 20, 20 ) }
      }
      DATASET "dset_03" {
         DATATYPE  H5T_IEEE_F64LE
         DATASPACE  SIMPLE { ( 1, 100, 20, 20 ) / ( 1, 100, 20, 20 ) }
      }
      DATASET "dset_04" {
         DATATYPE  H5T_IEEE_F64LE
         DATASPACE  SIMPLE { ( 1, 100, 20, 20 ) / ( 1, 100, 20, 20 ) }
      }
      DATASET "dset_05" {
         DATATYPE  H5T_IEEE_F64LE
         DATASPACE  SIMPLE { ( 1, 100, 20, 20 ) / ( 1, 100, 20, 20 ) }
      }    
   }
   GROUP "group2" {
   }
}
}

E:[.\StackOverflow]->
Related