I am trying to determine the fastest way to search a nested dictionary with a regular expression and to return the paths to each occurrence of that string. I am only interested in values that are strings, not other values that may not explicitly be strings. Recursion is not my strong suit. Here is an example JSON, let's say I'm looking for all absolute paths that contain 'blah'.
d = {'id': 'abcde',
'key1': 'blah',
'key2': 'blah blah',
'nestedlist': [{'id': 'qwerty',
'nestednestedlist': [{'id': 'xyz', 'keyA': 'blah blah blah'},
{'id': 'fghi', 'keyZ': 'blah blah blah'}],
'anothernestednestedlist': [{'id': 'asdf', 'keyQ': 'blah blah'},
{'id': 'yuiop', 'keyW': 'blah'}]}]}
I found the following code snippet, but am unsuccessful in making it return the paths, rather than just print them. Beyond that, it shouldn't be too hard to add an "IF value is a string AND contains re.search() THEN append the path to list".
def search_dict(v, prefix=''):
if isinstance(v, dict):
for k, v2 in v.items():
p2 = "{}['{}']".format(prefix, k)
search_dict(v2, p2)
elif isinstance(v, list):
for i, v2 in enumerate(v):
p2 = "{}[{}]".format(prefix, i)
search_dict(v2, p2)
else:
print('{} = {}'.format(prefix, repr(v)))