Bash - Multi-character string replacement when strings consist of unknown length but same character

Viewed 96

Assume a multi-line text string in which some lines start with a key-character ("#" in our case). Further assume that you wish to replace all instances of a target character ("o" in our case) with a different character ("O" in our case), if - and only if - that target character occurs as a string of two or more adjacent copies (e.g., "ooo"). This replacement is to be done in all lines that do not start with the key-character and must be case-sensitive.

For example, the following lines ...

#Foo bar
Foo bar
#Baz foo
Baz foo

are supposed to be converted into:

#Foo bar
FOO bar
#Baz foo
Baz fOO

The following attempt using sed does not retain the correct number of target characters:

$ echo -e "#Foo bar\nFoo bar\n#Baz foo\nBaz foo" | sed '/^#/!s/o\{2,\}/O/g'
#Foo bar
FO bar
#Baz foo
Baz fO

What code (with sed or otherwise) would conduct the desired replacement correctly?

3 Answers

Using sed

$ echo -e "#Foo bar\nFoo bar\n#Baz foo\nBaz foo" | sed '/#/!s/o\{2\}/\U&/'
#Foo bar
FOO bar
#Baz foo
Baz fOO

You can use Perl:

echo -e "#Foo bar\nFoo bar\n#Baz foo\nBaz foo" | perl -pe 's/^#.*(*SKIP)(*F)|o{2,}/"O" x length($&)/ge'

Here, ^#.*(*SKIP)(*F) matches and skips all lines starting with #, then o{2,} matches two or more o chars, and "O" x length($&) replaces these matches with O that is repeated the match size times ($& is the match value). Note the e flag after g that is used to evaluate the string on the right-hand side.

See the online demo:

#!/bin/bash
s="#Foo bar
Foo bar
#Baz foo
Baz foo"
perl -pe 's/^#.*(*SKIP)(*F)|o{2,}/"O" x length($&)/ge' <<< "$s"

Output:

#Foo bar
FOO bar
#Baz foo
Baz fOO

This might work for you (GNU sed):

sed -E '1{x;s/^/O/;x}
        /^#/b
       :a;/oo+/!b;s/oo+/\n&\n/;tb
       :b;G;s/\n\n(.*)\n.$/\1/;ta;s/\n[^\n](.*\n.*)\n(.)$/\2\n\1/;tb' file

In overview, use a designated character or characters (in this case O) to replace two or more o's where the start of a line is not #.

Prime the hold space with the designated character.

If the line starts with #, break out.

If the line does not contain two or o's, break out.

Otherwise, surround the two or more o's by newlines.

Append the replacement character and then replace non-newline characters between two newlines with the designated character.

When all replacements for the current set of o's have been replaced, check for more by continuing as above.

Once all replacements have been found, print the amended line.


A solution allowing for multiple replacements:

sed -E '1{x;s/^/oOxX/;x}
        /^#/b;
        :a;G;/((.)\2+)(.*\n(..)*\2)/!s/\n.*//;t;s//\n\1\n\3/;tb
        :b;s/\n\n(.*)\n.*$/\1/;ta;s/\n(.)(.*\n.*\n(..)*\1(.))/\4\n\2/;tb' file
Related