Match word only if no given prefix with arbitrary number of spaces

Viewed 97

I am trying to create a regex which would match the word bar except when there is the word foo behind.

I found that negative lookbehind could handle this, but the problem is that there is an arbitrary number of characters belonging to the expression [\s\-/] between foo and bar.

And unfortunately negative lookbehind does not support arbitrary length.

So the pattern (?<!foo[\s\-/]*)bar is not valid.

Do you know a regex technique that can overcome this problem?

4 Answers

One solution would be :

import re

c = re.compile(r'^(?!.*foo.*bar).*(bar).*$')

lst = ['bar', 'hi bar', 'foo   bar', 'foobar', 'hiiifoohiiibar']

for i in lst:
    match_obj = c.match(i)
    if match_obj:
        print(match_obj.group(), '|', match_obj.group(1))

output :

bar | bar
hi bar | bar

DEMO

Explanation : First we check the whole string to see if there is both 'foo' and 'bar' in the string(first foo then bar) with (?!.*foo.*bar). This is a negative lookahead assertion, if this pair doesn't exist, we can proceed.

Next that we are sure there isn't any foo before bar, we get all the string including bar. We put that in a group so that we can retrieve it via group(1).

One technique is to use the PyPi regex-module instead of the standard re-module. As I've read your query it looks like you would want to validate any string with the word "bar" in there unless it's preceded by the word "foo" along with an arbitrairy number of spaces and hyphens. If that's correct you can use:

(?<!foo[\s-]*)bar

Meaning; a negative lookbehind that starts with 'foo' and contains 0+ times a whitespace-character and/or hyphens. Here is some sample code:

import regex as re
lst = ['foobar', 'foo   -   bar', 'foo- -bar', 'foodbar']
for i in lst:
    if re.search(r'(?<!foo[\s-]*)bar', i):
        print(i)

Prints:

foodbar

You will need this pip package regex - it would not work with the default re:

foo\s*+bar(*SKIP)(*FAIL)|bar

regex101

An example invocation in the interpreter:

>>> import regex
>>> print(regex.search(r'foo\s*+bar(*SKIP)(*FAIL)|bar', 'fdfdf foo bar fdfdf foo bar bar'))
<regex.Match object; span=(28, 31), match='bar'>

My solution is simple: There are two parts to the test:

  1. If "bar" is in the text
  2. If not ("bar", plus [\s\-/], plus "foo")

Putting it into code:

import re

data = [
    # Good
    "bar and not foo",
    "bar alone",

    # Bad
    "bar - foo",
    "barfoo",
    "bar foo",
    "bar / foo",
]


for text in data:
    if "bar" in text and not re.match(r"bar[\s\-/]*foo", text):
        print(text)

Output:

bar and not foo
bar alone

In general, I stay away from regular expression because it is hard to understand. I only use it when I must.

Related