I agree with the other commenters that there are much better ways to develop a parser like this.
Nevertheless, I want to suggest two proposals:
while(/\G((?:[^;'"\\]++|'[^']*+'|"[^"]*+"|\\.)*;)/gx){
print " command: $1\n";
# process the match . . .
}
\G is a zero-width assertion that matches the position where the previous m//g left off, see perlop#\G-assertion. (In the docs there is also an example of a lex-like scanner that might be of interest.)
- The non-capturing group contains harmless chars, quoted strings, and escaped characters
- Note the use of possessive quantifiers in order to avoid performance issues due to backtracing.
- i removed the negative assertion
(?<!\\), because this would fail in cases such as echo \\;
This code will work with your given examples. However, e.g. bash allows escaping double-quotes inside of double-quoted string such as echo "\"".
If your shell should accept such a code, too, then the regexp has to be expanded:
while(/\G( # anchor for beginning
(?:[^;'"\\]++ # harmless chars
|'[^']*+' # or single-quoted string
|"(?: # or double-quoted string,
[^"\\]++ # containing harmless chars
|\\. # or an escaped char
)*+" # with arbitrary many repetitions
|\\. # or an escaped char
)*+ # with arbitrary many repetitions
;) # end with semi-colon
/gx){
print " command: $1\n";
# process the match . . .
}
Such pure regexp solutions are very error-prone. And the more exceptions you find that have to be treated, the more complicated the pattern get and the more difficult it gets to debug that code.
some tests:
use strict;
use warnings;
use Test::More tests => 16;
my $samples = [
{"'some ; text ;'" => []},
{'echo;' => ['echo;']},
{'echo ";ignore; inside ;" ; echo \'something;\' \; \'else\';' => [
'echo ";ignore; inside ;" ;', ' echo \'something;\' \; \'else\';']},
{'echo moep; echo moep;' => [ 'echo moep;', ' echo moep;']},
{'echo \a ; echo moep;' => [ 'echo \a ;', ' echo moep;']},
{'echo \\a ; echo moep;' => [ 'echo \\a ;', ' echo moep;']},
{'echo \\\a ; echo moep;' => [ 'echo \\\a ;', ' echo moep;']},
{'echo \; echo moep;' => [ 'echo \; echo moep;']},
{'echo \\; echo moep;' => [ 'echo \; echo moep;']}, # '\\;' eq '\;' !
{'echo \\\; echo moep;' => [ 'echo \\\;', ' echo moep;']},
{'echo ";\';\';"; echo moep;' => [ 'echo ";\';\';";', ' echo moep;']},
{'echo "\";"; echo moep;' => [ 'echo "\";";', ' echo moep;']},
{'echo ";\""; echo moep;' => [ 'echo ";\"";', ' echo moep;']},
{'echo "\";\""; echo moep;' => [ 'echo "\";\"";', ' echo moep;']},
{'echo ";\\\\"; echo moep;' => [ 'echo ";\\\\";', ' echo moep;']},
{'echo "\\\\\";\""; echo moep;' => [ 'echo "\\\\\";\"";', ' echo moep;']},
];
for my $sample(@$samples){
while(my ($line, $test) = each %$sample){
my @result = $line =~ /\G((?:[^;'"\\]++|'[^']*+'|"(?:[^"\\]++|\\.)*+"|\\.)*+;)/g;
is_deeply(\@result, $test, $line);
}
}
Still, you can easily find many false positive/negative samples. For example I did not cope with parentheses. This would make the above pattern much more complicated by using recursive subpatterns.