Can re ignore a lazy quantifier?

Viewed 82

Given this code (Python 3.6):

>>> import re
>>> a = re.search(r'\(.+?\)$', '(canary) (wharf)')
>>> a
<_sre.SRE_Match object; span=(0, 16), match='(canary) (wharf)'>
>>>

Why doesn't re stop searching at the first parethesis closure?

The expected output is None. The search should detect that there is not an end of line after (canary), but it doesn't.

Edit:If there is only ONE word between parens, it should match, if there are more than one, it shouldn't match at all.

Any help would be hugely appreciated.

2 Answers

The lazy flag isn't being ignored.

You get a match on the entire string because .+? means match anything one or more times until you find a match, expanding as needed. If the regex was \([^)]+?\)$ it would have matched only the last (wharf) because we excluded the +? from matching )

Or if the regex was \(.+?\), it would have matched the (canary) and the (wharf), which shows that it's being lazy.

\(.+?\)$ matches everything because you make it match everything until the end of the line.

If you want to ensure that there is only one group in parentheses in the entire string, we can do that with our "no-parentheses-regex" from above and force the start of the string to match the start of your regex.

^\([^)]+?\)$
Try it: https://regex101.com/r/Ts9JeF/1

Explanation:

  • ^\(: Match a literal ( at the start of the string
  • [^)]+?: Match anything but ), as many times as needed
  • \)$: Match a literal )$ at the end of the line.

Or, if you want to allow other words before and after the one in parentheses, but nothing in parentheses, do this:

^[^()]*?\([^)]+?\)[^()]*$
Try it: https://regex101.com/r/Ts9JeF/3

Explanation:

  • ^[^()]*?: At the start of the string, match anything but parentheses zero or more times.
  • \([^)]+?\): Very similar to our previous regex
  • [^()]*$: Match zero or more non-parentheses characters until the end of the string.

the non-greedy qualifier makes it match the shortest repeat -- in this case the shortest successful repeat is the entire string. it doesn't "not match the )" because you didn't tell it to do so

you can think of the engine doing something like this (using simplified string '(a) (b)':

  1. start at position 0
  2. '(' matches (, proceed to position 1
  3. 'a' matches ., proceed to position 2
    • (non-greedy) ')' matches ), proceed to position 3
    • (non-greedy) end of string does not match $ => backtrack to position 2
  4. ')' matches . proceed to position 3
    • (non-greedy) ' ' does not match )
  5. ' ' matches . proceed to position 4
    • (non-greedy) '(' does not match )
  6. '(' matches . proceed to position 5
    • (non-greedy) 'b' does not match )
  7. 'b' matches . proceed to position 6
    • (non-greedy) ')' matches )
    • (non-greedy) $ matches end of string => DONE!

try this regex on for size:

r'\([^)]+\)$'

here a left-paren is matched, followed by a nonzero number of non-right parens followed by a right paren and the end of the string

Related