POSIX: abcdef to ab bc cd de ef

Viewed 264

Using POSIX sed or awk, I would like to duplicate every second character in every pair of neighboring characters and list every newly-formed pair on a new line.

example.txt:

abcd 10001.

Expected result:

ab
bc
cd
d 
 1
10
00
00
01
1.

So far, this is what I have (N.B. omit "--posix" if on macOS). For some reason, adding a literal newline character before \2 does not produce the expected result. Removing the first group and using \1 has the same effect. What am I missing?

sed --posix -E -e 's/(.)(.)/&\2\
/g' example.txt

abb
cdd
100
000
1..
7 Answers

Try:

$ echo "abcd 10001." | awk '{for(i=1;i<length($0);i++) print substr($0,i,2)}'
ab
bc
cd
d 
 1
10
00
00
01
1.

You may use

sed --posix -e 's/./&\
&/g' example.txt | sed '1d;$d'

The first sed command finds every char in the string and replaces with the same char, then a newline and then the same char again. Since it replaces first and last chars, the first and last resulting lines must be removed, which is achieved with sed '1d;$d'.

Had sed supported lookarounds, one could have used (?!^).(?!$) (any char but not at the start or end of string) and the last sed command would not have been necessary, but it is not possible with sed. You could use it in perl though, perl -pe 's/(?!^).(?!$)/$&\n$&/g' example.txt (see demo online, $& in the RHS is the same as & placeholder in sed, the whole match value).

With GNU awk could you please try following. Written and tested with shown samples and tested it in link https://ideone.com/qahp0S

awk '
BEGIN{
  FS=""
}
{
  for(i=1;i<=(NF-1);i++){
    print $i$(i+1)
  }
}
' Input_file

Explanation: setting field separator as NULL in the BEGIN section of program for all lines here. Then in main program running a for loop which runs from 1st field to till 2nd last field. In that loop's each iteration printing current and next field.

Using same routine, it can be done in bash itself:

s='abcd 10001.'

for((i=0; i<${#s}-1; i++)); do echo "${s:i:2}"; done
ab
bc
cd
d
 1
10
00
00
01
1.

Just for fun, a single sed consisting of 3 substitutions:

$ echo "abcd 10001." | sed 's/./&&/g;s/\(^.\|.$\)//g;s/../&\n/g'

The first part duplicates all characters, the second part removes the first and last character, the third part adds a newline character after each character-pair.

If you want to be POSIX compliant you have to do:

$ echo "abcd 10001." | sed -e  's/./&&/g' -e 's/^.//g' -e 's/.$//g' -e 's/../&\n/g'

Here we had to add an extra one as the expression \(^.\|.$) is an ERE and posix sed only accepts a BRE

Process substitution isn't specified by POSIX. The POSIX requirement was only specified for awk and sed, so maybe the next solution is acceptable:

paste -d '\0' <(echo; fold -w1 example.txt) <(fold -w1 example.txt) | grep ..

or

while read -n1 ch; do
   printf "%s\n%s" "${ch}" "${ch}"
done < example.txt | grep ..

or

sed 's/./&&/g;s/.//' example.txt | grep -o ..

This might work for you (GNU sed):

sed 's/.\(.\)/&\n\1/;/../P;D' file

Replace the first two characters by the first two characters, a newline and the second character.

Print the first line if it is two characters long, delete the first line and repeat.

Alternative, more long winded:

sed -E ':a;s/^(([^\n]{2}\n)*[^\n])([^\n])([^\n])/\1\3\n\3\4/;ta' file

Or, with no hardcoded new line:

sed -E '/.../{G;s/^(.(.))(.*)(.)/\1\4\2\3/;P;D}' file

Lastly:

sed 's/./&\n&/g;s/^..\|..$/g' file
Related