parsing a string with python regex optional named groups

Viewed 143

I am struggling with python named group re

I have the following string: "blah blah id=xyz, blah blah foo bar=zxy] blah baz=a}" terminating chars are: ,]}

and I would like to get out named dict using regex pattern that looks like this:

{'id': 'xyz', 'foo bar': 'zxy', 'baz': 'a'} groups should be optional

I was able to hit it without named groups and including termination characters but I am sure there is a way how to do it fully in regexp and be more elegant ... it just eludes me any help would be welcome

my current solution is using folowing pregmatch:

(id=.* ?[, }\]] |baz=.* ?[, }\]] |foo bar=.* ?[, }\]])

it works but requires significant postprocessing (string splitting and striping)

         for i in ae2:
            key, value = i.split('=', 1)
            altevent2[key] = value.strip('},] ')

Also, it would be cool to get rid of whitespace/unprintable chars but only when they are at the start/end of the value

if at all possible it should require no postprocessing - I need a lot of performance

Edit1: list if dict 'IDs' is known in advance, for this case it would be 'id','foo bar','baz'

2 Answers

You can use the re.split() method to split the initial string from your endpoints and search for matching with your keys in your dict, like the code.

import re
    
ex = "blah blah id=xyz, blah blah foo bar=zxy] blah baz=a}"
dict_keys = ["id", "foo bar", "baz"]

end = re.split(", |] |}", ex)  # ['blah blah id=xyz', ' blah blah foo bar=zxy', ' blah baz=a', '']

result = {}

for i in dict_keys:
    for j in end:
        if i in j:
            result[i] = j.partition("=")[2]

OBs: I try to avoid the "2 for", but I could't find a way to do that.

One simple solution is using re.findall.

s = "blah blah id=xyz, blah blah foo bar=zxy] blah baz=a}"
re.findall('(id|foo bar|baz)=([^,}\]]+)', s)
Related