Subclassing pandas Index object

Viewed 58

I am trying to create a subclass to a pandas Index object that does simple type checking and enforces naming conventions. My application has a set of physical machines which measure data, each with its own numeric identifier. I have a set of measurements from different machines and want to keep track of which machine data was collected on. I want my subclass to check that the machine number is valid and also inherit pd.Index behavior and methods, like loc.

When I try to subclass pd.Index I get unexpected behavior. First, it tells me the type of my object is pandas.core.indexes.numeric.Int64Index, not my subclass:

import pandas as pd

class MachineIndex(pd.Index):
    MACHINE_IDs = [7, 22, 24]    # valid machine ids
    def __init__(self, ids):
        assert all([_id in MachineIndex.MACHINE_IDs for _id in ids]), "invalid id"
        super().__init__(data=ids, name='Machine ID')
        self._hi = 'hi'

I = MachineIndex([22, 7, 7, 22, 24, 22])
print(type(I))  # pandas.core.indexes.numeric.Int64Index  instead of MachineIndex

Second, I am not able to access a dummy attribute that I put on the object:

print(I._hi)    # AttributeError: 'Int64Index' object has no attribute '_hi'

Finally, I run into difficulties specifying the parameter name to the constructor:

I = MachineIndex(ids=[22, 7, 7, 22, 24, 22])  # TypeError: MachineIndex(...) must be called with a collection of some kind, None was passed

Is there a way to fix these errors and is subclassing the best approach for Index objects? I notice the pandas docs suggests registering custom accessors rather than subclassing for DataFrame objects. Does a similar philosophy apply to Index objects and if so is there example code available?

0 Answers
Related