Perl regex for a character NOT within string characters

Viewed 132

I am writing a perl script that 'compiles' shell code. One thing I need to do is detect ; characters and deal with them (things like multiple commands on one line), but only when they are not escaped (by \ ), or within a string. For example, we shouldn't match 'some ; text ;' , but we should match the semicolons in between the two echo statements in echo ";ignore; inside ;" ; echo 'something;' \; 'else';

In the above example, exactly TWO semicolons should have been matched.

I have tried this with a regex loop

while ($_ =~ /('[^']+')*?("[^"]+")*?(?<!\\)(?<match>;)/g) 
  { 
    print "semiolon: $+{match}\n"; 
    # process the match . . . 
  }

Whilst this works for some examples, there are some cases where it doesn't properly detect the semicolon is 'inside' two strings; as it can't match a PAIR of them before the current match. How would I go about ensuring that we only match semicolons outside a string?

Thanks in advance.

1 Answers

I agree with the other commenters that there are much better ways to develop a parser like this.

Nevertheless, I want to suggest two proposals:

while(/\G((?:[^;'"\\]++|'[^']*+'|"[^"]*+"|\\.)*;)/gx){ 
    print "  command: $1\n"; 
    # process the match . . . 
}
  • \G is a zero-width assertion that matches the position where the previous m//g left off, see perlop#\G-assertion. (In the docs there is also an example of a lex-like scanner that might be of interest.)
  • The non-capturing group contains harmless chars, quoted strings, and escaped characters
  • Note the use of possessive quantifiers in order to avoid performance issues due to backtracing.
  • i removed the negative assertion (?<!\\), because this would fail in cases such as echo \\;

This code will work with your given examples. However, e.g. bash allows escaping double-quotes inside of double-quoted string such as echo "\"". If your shell should accept such a code, too, then the regexp has to be expanded:

while(/\G( # anchor for beginning
    (?:[^;'"\\]++            # harmless chars
    |'[^']*+'                # or single-quoted string
    |"(?:                    # or double-quoted string,
        [^"\\]++               #   containing harmless chars
        |\\.                   #   or an escaped char
        )*+"                   #   with arbitrary many repetitions
    |\\.                     # or an escaped char
    )*+                      # with arbitrary many repetitions
    ;)                       # end with semi-colon
/gx){ 
    print "  command: $1\n"; 
    # process the match . . . 
}

Such pure regexp solutions are very error-prone. And the more exceptions you find that have to be treated, the more complicated the pattern get and the more difficult it gets to debug that code.

some tests:

use strict;
use warnings;
use Test::More tests => 16;

my $samples = [
    {"'some ; text ;'" => []},
    {'echo;' => ['echo;']},
    {'echo ";ignore; inside ;" ; echo \'something;\' \; \'else\';' => [
            'echo ";ignore; inside ;" ;', ' echo \'something;\' \; \'else\';']},
    {'echo moep; echo moep;' => [ 'echo moep;', ' echo moep;']},
    {'echo \a ; echo moep;' => [ 'echo \a ;', ' echo moep;']},
    {'echo \\a ; echo moep;' => [ 'echo \\a ;', ' echo moep;']},
    {'echo \\\a ; echo moep;' => [ 'echo \\\a ;', ' echo moep;']},
    {'echo \; echo moep;' => [ 'echo \; echo moep;']},
    {'echo \\; echo moep;' => [ 'echo \; echo moep;']}, # '\\;' eq '\;' !
    {'echo \\\; echo moep;' => [ 'echo \\\;', ' echo moep;']},
    {'echo ";\';\';"; echo moep;' => [ 'echo ";\';\';";', ' echo moep;']},
    {'echo "\";"; echo moep;' => [ 'echo "\";";', ' echo moep;']},
    {'echo ";\""; echo moep;' => [ 'echo ";\"";', ' echo moep;']},
    {'echo "\";\""; echo moep;' => [ 'echo "\";\"";', ' echo moep;']},
    {'echo ";\\\\"; echo moep;' => [ 'echo ";\\\\";', ' echo moep;']},
    {'echo "\\\\\";\""; echo moep;' => [ 'echo "\\\\\";\"";', ' echo moep;']},
];

for my $sample(@$samples){
    while(my ($line, $test) = each %$sample){
        my @result = $line =~ /\G((?:[^;'"\\]++|'[^']*+'|"(?:[^"\\]++|\\.)*+"|\\.)*+;)/g;
        is_deeply(\@result, $test, $line);
    }
}

Still, you can easily find many false positive/negative samples. For example I did not cope with parentheses. This would make the above pattern much more complicated by using recursive subpatterns.

Related