Using Python 'not in' on dict with tuple key where I don't have all tuple parts

Viewed 267

I have a script that reads a list of users and reports from a MySQL database and runs a function for each of the values. The script checks the db every 5 minutes and adds these value pairs to a dictionary. As I loop through I don't want to re-add value pairs that are still in the dictionary. I've had no issues handling that with this code:

if (user, report) not in t1.reports_to_call:
    t1.add(report, user)

My issue is that I want to also check a dictionary t2.reports_requested where the key is a tuple of three parts user, report and some_unknown_id. If there a way to run not in on a dictionary where you only know 2 of three elements and want to wildcard the third?

I have been also looking to see if I can refactor this secondary dictionary and move some_unknown_id out of the key as I do think it may be the case that the user/report pair is unique in this dictionary. But if I determine I need to wildcard the third value what is best way to do this?

{('user', 'report', 'some_unknown_id'): {}}
5 Answers

You can use set and check if a new pair it is a subset of the t2.reports_requested.keys()

values_exist = [True for k in reports_requested.keys() if set(["user","report"]).issubset(k)]
if not values_exist:
    print("Add new tuple")
    pass

For example:

reports_requested = {('user', 'report', 'some_unknown_id'): {}, ('user1', 'report', 'some_unknown_id'): {}, ('user1', 'report', 'some_unknown_id2'): {}}
new_pair = {'user1', 'report2'}

values_exist = [True for k in reports_requested.keys() if new_pair.issubset(k)]
if not values_exist:
#     print("Add new tuple")
    pass

Performance check with timeit: 1.06 µs ± 184 ns per loop (mean ± std. dev. of 7 runs, 1000000 loops each)

No, you cannot check for the existence of keys in a dict using wild-cards.

Make a dict of dicts:

{('user', 'report'): {
    'some_unknown_id': {},
}

or even better:

{'user': {
    'report': {
        'some_unknown_id' : {},
    },
}

Then you have the best of all worlds.

You can first check whether a user occurs, and if so, whether in combination with a report, and if so, whether in combination with some other identifier.

You can use zip to extract "sub-keys" for testing against, as

if (user,report) not in zip(*list(zip(*t2.keys()))[0:2]):
  #etc

Here, the [0:2] extracts only the first two components of the key.

You can't wildcard the key but you could replace the third key with None if it isn't there:

x = {('user','report'): {'dummy': 1}, ('user','report', 'unknown_key'): {'dummy': 2}, }

new_dict = dict(zip({(k[0], k[1], k[2] if len(k) == 3 else None) for k in x.keys()}, x.values()))

new_dict

{('user', 'report', 'unknown_key'): {'dummy': 1},
 ('user', 'report', None): {'dummy': 2}}

This way you'll always know the length of the key and you can test the last value.

You can rely on the default get operation of the dictionary and add such element if None is returned as follow.

if not t1.get((user, report), default=None):
    t1.add(report, user)

for the t2, I'd suggest to do a two level dictionary, sort of:

t2 = {
    ('user', 'report'): {
        'some_unknown_id': {},
        'some_unknown_id_2': {},
    }
}

In that fashion you can handle several "third unkown keys" right under ('user', 'report').

Related