How to negate the sequences of characters in the positive lookbehind using regexes?

Viewed 72

Given the following strings:

<i>aaa>
<i>aaa>>
<i>AAA>
<i>AAA>>
<i>999>
<i>9>
<i>>
<b>aaa>
<b>AAA>

I want to use regular expressions to match one or more final angle brackets > if the string contains <i> followed by some sequence of characters.

I have tried using the positive lookbehind: (?<=<i>[A-Za-z\d].*)>.* in order to ignore the <i> and some sequence of characters until the final bracket, but got the error * A quantifier inside a lookbehind makes it non-fixed width.

How to group the characters inside the positive lookbehind?

1 Answers

You may use

re.sub(r'(<i>[A-Za-z\d]*)>+$', r'\1</i>', text)

Or, a bit more generic:

re.sub(r'(<i>.*?)>+$', r'\1</i>', text)   # if there can be anything after <i>
re.sub(r'(<i>[^>]*)>+$', r'\1</i>', text) # if there can be anything but > after <i>

Or even

re.sub(r'(<i>[^>]*)>+$', r'\1</i>', text, flags=re.M) # To replace at each line end

See regex demo.

Pattern details

  • (<i>[A-Za-z\d]*) - a capturing group that matches and places in Group 1 (its value is referred to with \1 from the replacement pattern) <i> and then 0 or more ASCII letters and digits
  • [^>]* - matches 0 or more chars other than >
  • .*? - matches 0 or more chars other than line break chars, as few as possible
  • >+ - 1 or more > chars
  • $ - end of string (or line if re.M flag is provided).
Related