pickle: slow dict deserialization

Viewed 503

Let's suppose I have a "quite large" dictionary, where keys are objects with heavy __eq__ function

class MyObject():

     def __eq__(self, other):
         return <very heavy function call>

     def __hash__(self):
         return <not so heavy hash calculation>

mydict = {<MyObject>:<Int>}

The problem is that when I try to unpickle a pickled object, it requires much time. I believe that it is because pickle does not save the internal hash table, and when restoring it recalculates the dictionary.

I did a simple experiment:

import pickle 
import pickletools 
original = { 'a': 0, 'b': [1, 2, 3] } 
pickled = pickle.dumps(original) 
pickletools.dis(pickled)

results in:

    0: \x80 PROTO      3
    2: }    EMPTY_DICT
    3: q    BINPUT     0
    5: (    MARK
    6: X        BINUNICODE 'a'
   12: q        BINPUT     1
   14: K        BININT1    0
   16: X        BINUNICODE 'b'
   22: q        BINPUT     2
   24: ]        EMPTY_LIST
   25: q        BINPUT     3
   27: (        MARK
   28: K            BININT1    1
   30: K            BININT1    2
   32: K            BININT1    3
   34: e            APPENDS    (MARK at 27)
   35: u        SETITEMS   (MARK at 5)
   36: .    STOP
highest protocol among opcodes = 2

There are no signs of a hashtree. It means that Pickle Machine has to recalculate hashtree after deserialization. But it is really dummy? Why pickle does not save the internal state of a dictionary and how one can struggle with it?

3 Answers
  • The __eq__ method is only called on insertion into a dict when there's a hash collision (the __hash__ method for different objects returns the same answer); if this is happening often enough to be a noticeable slow-down, it'll also massively slow down all other operations — in the extreme, if the __hash__ method always returns the same answer regardless of the value of the object, the __eq__ method will be called ½n² times just for the unpickle.

    If this is happening, you need to improve your __hash__ method so that it returns better answers (with fewer duplicates for different values of your object). You can check by manually collecting the results of the __hash__ method on a representative sample of your objects and making sure that they're mostly different.

  • If the above doesn't resolve the issue, I would suggest to profile the code to confirm where the bottleneck is; then you can see if you can speed up the function that's the bottleneck.

The Pickle Machine is implemented across various Python Objects, and certain objects may not like a hash implementation, and may brake.

If, all your going do is save Json Objects, use the json object as its will be efficient.

Use pickled objects as dictionary keys instead.

By using the pickled object as the key, we can circumvent the MyObject.__hash__ and MyObject.__eq__ methods altogether and potentially speed up the unpickling process.

Once you have the unpickled dictionary, you can slowly change the pickeled keys to actual MyObject instances as you need them.

p1.py

Create a dictionary where keys are of type MyObject then pickle them to a file. Create a dictionary where the keys are pickled MyObject bytes. Compare the times required to unpickle each dictionary.

#!/usr/bin/env python3

import pickle
import pickletools
from pprint import pprint
import random
import string
import time

random.seed(19891225)


class MyObject(object):

    def __init__(self, value):
        self.value = value

    def __eq__(self, other):
        exit(1)
        return self.value == other.value

    def __hash__(self):
        time.sleep(10)
        return hash(tuple(self.value))


def key_value():
    value_len = random.randint(0,len(string.ascii_letters))
    value = random.sample(string.ascii_letters, value_len)
    return (MyObject(value), value_len)

def object_key_dict(size=10):
    return dict([key_value() for i in range(0, size)])

def pickle_key_dict(d):
    return {pickle.dumps(k):v for k, v in d.items()}

def pickle_unpickle(obj, filename):
    with open(filename, "wb") as pout:
        pickle.dump(obj, pout)

    print(f"Unpickle Start: {time.strftime('%H:%M:%s')}")
    with open(filename, "rb") as pin:
        newobj = pickle.load(pin)
    print(f"Unpickle complete: {time.strftime('%H:%M:%s')}")
    print("-"*20)

    return newobj


if __name__ == "__main__":
    d = object_key_dict(size=10)

    print("Use MyObject instances as dictionary keys")
    unpickled_d = pickle_unpickle(d, 'doc.p')

    print("\nUse pickled ojbects as dicitonary keys")
    pd = pickle_key_dict(unpickled_d)
    upd = pickle_unpickle(pd, 'docp.p')

Output

Use MyObject instances as dictionary keys
Unpickle Start: 09:25:1589808319
Unpickle complete: 09:26:1589808419
--------------------

Use pickled ojbects as dicitonary keys
Unpickle Start: 09:26:1589808419
Unpickle complete: 09:26:1589808419
--------------------
Related