Nested dataclass initialization

Viewed 807

I have a JSON object that reads:

j = {"id": 1, "label": "x"}

I have two types:

class BaseModel:
    def __init__(self, uuid):
        self.uuid = uuid

class Entity(BaseModel):
    def __init__(self, id, label):
        super().__init__(id)
        self.name = name

Note how id is stored as uuid in the BaseModel.

I can load Entity from the JSON object as:

entity = Entity(**j)

I want to re-write my model leveraging dataclass:

@dataclass
class BaseModel:
    uuid = str

@dataclass
class Entity:
    name = str

Since my JSON object does not have the uuid, entity = Entitye(**j) on the dataclass-based model will throw the following error:

TypeError: __init__() got an unexpected keyword argument 'id'

The "ugly" solutions I can think of:

  • Rename id to uuid in JSON before initialization:

    j["uuid"] = j.pop("id")
    
  • Define both id and uuid:

    @dataclass 
    class BaseModel:
        uuid = str
    
    @dataclass
    class Entity:
        id = str
        name = str
    
        # either use:
        uuid = id
        # or use this method
        def __post_init__(self):
            super().uuid = id
    

Is there any cleaner solution for this kind of object initialization in the dataclass realm?

3 Answers

might be ruining the idea of removing the original __init__ but how about writing a function to initialize the data class?

def init_entity(j):
    j["uuid"] = j.pop("id")
    return Entity(**j)

and in your code entity = initEntity(j)

I think the answer here might be to define a classmethod that acts as an alternative constructor to the dataclass.

from dataclasses import dataclass
from typing import TypeVar, Any

@dataclass
class BaseModel:
    uuid: str


E = TypeVar('E', bound='Entity')


@dataclass
class Entity(BaseModel):
    name: str

    @classmethod
    def from_json(cls: type[E], **kwargs: Any) -> E:
        return cls(kwargs['id'], kwargs['label']

(For the from_json type annotation, you'll need to use typing.Type[E] instead of type[E] if you're on python <= 3.8.)

Note that you need to use colons for your type-annotations within the main body of a dataclass, rather than the = operator, as you were doing.

Example usage in the interactive REPL:

>>> my_json_dict = {'id': 1, 'label': 'x'}
>>> Entity.from_json(**my_json_dict)
Entity(uuid=1, name='x')

It's again questionable how much boilerplate code this saves, however. If you find yourself doing this much work to replicate the behaviour of a non-dataclass class, it's often better just to use a non-dataclass class. Dataclasses are not the perfect solution to every problem, nor do they try to be.

Simplest solution seems to be to use an efficient JSON serialization library that supports key remappings. There are actually tons of them that support this, but dataclass-wizard is one example of a (newer) library that supports this particular use case.

Here's an approach using an alias to dataclasses.field() which should be IDE friendly enough:

from dataclasses import dataclass

from dataclass_wizard import json_field, fromdict, asdict


@dataclass
class BaseModel:
    uuid: int = json_field('id', all=True)


@dataclass
class Entity(BaseModel):
    name: str = json_field('label', all=True)


j = {"id": 1, "label": "x"}

# De-serialize the dictionary object into an `Entity` instance.
e = fromdict(Entity, j)

repr(e)
# Entity(uuid=1, name='x')

# Assert we get the same object when serializing the instance back to a
# JSON-serializable dict.
assert asdict(e) == j
Related