I am trying to create a subclass to a pandas Index object that does simple type checking and enforces naming conventions. My application has a set of physical machines which measure data, each with its own numeric identifier. I have a set of measurements from different machines and want to keep track of which machine data was collected on. I want my subclass to check that the machine number is valid and also inherit pd.Index behavior and methods, like loc.
When I try to subclass pd.Index I get unexpected behavior. First, it tells me the type of my object is pandas.core.indexes.numeric.Int64Index, not my subclass:
import pandas as pd
class MachineIndex(pd.Index):
MACHINE_IDs = [7, 22, 24] # valid machine ids
def __init__(self, ids):
assert all([_id in MachineIndex.MACHINE_IDs for _id in ids]), "invalid id"
super().__init__(data=ids, name='Machine ID')
self._hi = 'hi'
I = MachineIndex([22, 7, 7, 22, 24, 22])
print(type(I)) # pandas.core.indexes.numeric.Int64Index instead of MachineIndex
Second, I am not able to access a dummy attribute that I put on the object:
print(I._hi) # AttributeError: 'Int64Index' object has no attribute '_hi'
Finally, I run into difficulties specifying the parameter name to the constructor:
I = MachineIndex(ids=[22, 7, 7, 22, 24, 22]) # TypeError: MachineIndex(...) must be called with a collection of some kind, None was passed
Is there a way to fix these errors and is subclassing the best approach for Index objects? I notice the pandas docs suggests registering custom accessors rather than subclassing for DataFrame objects. Does a similar philosophy apply to Index objects and if so is there example code available?