suppress empty string in pyparsing

Viewed 454

I am trying to parse a string data with nested curly braces, and I would like to remove the quotation marks from the parsed string elements, also ignore ",", "+" and empty string "" from the parsed list.

Here is what I have for the parser definition using pyparsing:

data = '{"url", {action1, action2}, "", class1 + class2}'
quotedString.setParseAction(removeQuotes)
parser = nestedExpr(opener="{", closer="}", ignoreExpr=(quotedString | Suppress(",") | Suppress("+")))
print(parser.parseString(data, parseAll=True)[0])

Here is the output: ['url', ['action1', 'action2'], '', 'class1', 'class2']

I would like to ignore the empty string '' during parsing and in the output, I tried adding Suppress(",") in ignoreExpr, but it seems the program goes into loop and give warning SyntaxWarning: null string passed

Also, is there a way to combine the Suppress() strings into a list instead of writing them one by one?

Thanks in advance.

1 Answers

nestedExpr is really just a sort of "cheater" expression, for easy skipping over of nested lists in parentheses, braces, etc. To actually parse the contents, or do meaningful processing of the contents, it is clearer to define an actual recursive expression (though this is some extra work).

I was not above some level of cheating myself though. I defined word_sum as a delimited list of words with '+' separators. This takes care of suppressing the '+' signs. I then used delimitedList again for the ',' separated parts, which again just give back a list of the list items, and drops the delimiting ','s. With these shortcuts, the recursive grammar ends up looking pretty brief. See the comments in the annotated code below.

(To your question about suppressing multiple characters with listing them all, you can do this using pyparsing oneOf: pp.Suppress(pp.oneOf('+ - , *')) or pp.oneOf("+ - , *").suppress(), whichever suits your taste.

import pyparsing as pp

# delimitedList will suppress the delimiters, and just return the list elements
# delimitedList also will match just a single word
word = pp.Word(pp.alphas, pp.alphanums)
word_sum = pp.delimitedList(word, delim="+")

# expression for items that can be found inside the braces, in a list delimited by commas
# - define an explicit suppressor for ""
# - match QuotedStrings
# - match word_sums
item = pp.Literal('""').suppress() | pp.QuotedString('"') | word_sum

# define a Forward for the recursive expression
brace_expr = pp.Forward()

# define the contents of the recursive expression, which can include a reference to itself
# (use '<<=', not '=' for this definition)
LBRACE, RBRACE = map(pp.Suppress, "{}")
brace_expr <<=pp.Group(LBRACE + pp.delimitedList(item | brace_expr, delim=",") + RBRACE)

# try it out!
text = data = '{"url", {action1, action2}, "", class1 + class2}'
print(brace_expr.parseString(text)[0])

# prints
# ['url', ['action1', 'action2'], 'class1', 'class2']
Related