Apply substitution N times

Viewed 94

I have a string like

data1_data2_data3_data4@data5,data6

It happens that, sometimes, data5 contains underscores, which happens to be the field separator. Ugly, I know.

I want to read this data pieces with something like:

IFS="_@," read d1 d2 d3 d4 d5 d6 <<< "$input"

The problem arrives when data5 contains an underscore. To work around this issue. I want to substitute the first three underscores with commas (and the @ too). The easies way I found so far is with sed:

sed 's/_/,/; s/_/,/; s/_/,/; s/@/,/' <<< "$input"

But repeating three times the same substitution seems quite inefficient. What happens if I need to repeat it 5000 times?

Is there any way to tell sed to repeat a substitution a certain amount of times?

To be complete, sample input:

input="data1_data2_data3_data4@d_a_t_a_5,data6"
IFS="," read d1 d2 d3 d4 d5 d6 <<< "$input"

Expected output:

d1=="data1"
d2=="data2"
d3=="data3"
d4=="data4"
d5=="d_a_t_a_5"
d6=="data6"
5 Answers

You may use this awk in a process substitution:

input="data1_data2_data3_data4@d_a_t_a_5,data6"

IFS=, read d1 d2 d3 d4 d5 d6 < <(awk -F@ -v OFS=, -v n=3 '{
while (i++<n) sub(/_/, ",", $1)} 1' <<< "$input")

# check variable values
declare -p d1 d2 d3 d4 d5 d6

declare -- d1="data1"
declare -- d2="data2"
declare -- d3="data3"
declare -- d4="data4"
declare -- d5="d_a_t_a_5"
declare -- d6="data6"
  • awk command uses @ as field separator.
  • awk command replaces _ with , in 1st field only and exactly n times.

Use awk.

$ input="data1_data2_data3_data4@d_a_t_a_5,data6"
$ awk -v RS='[@\n]' '{ if(NR % 2){ gsub(/_/, ","); ORS = "," } else ORS = "\n" } 1' <<< "$input"
data1,data2,data3,data4,d_a_t_a_5,data6

an option could be to split manually using shell expansion ${var%%pat} strips largest suffix matching pat and ${var#pat} strips shortest prefix matching pat

while IFS= read line; do
    tmpline=$line
    d1=${tmpline%%_*} tmpline=${tmpline#*_}
    d2=${tmpline%%_*} tmpline=${tmpline#*_}
    d3=${tmpline%%_*} tmpline=${tmpline#*_}
    d4=${tmpline%%@*} tmpline=${tmpline#*@}
    d5=${tmpline%%,*} tmpline=${tmpline#*,}
    d6=${tmpline}

    printf "%s\n" "d1=$d1" "d2=$d2" "d3=$d3" "d4=$d4" "d5=$d5" "d6=$d6"
done <<< "$input"

or to get around bash read slowness, split lines manually

tmpinput=$input
while [[ $tmpinput ]]; do
    if [[ $tmpinput = *$'\n'* ]]; then
        tmpline=${tmpinput%%$'\n'*} tmpinput=${tmpinput#*$'\n'}
    else
        tmpline=${tmpinput} tmpinput=''
    fi

    d1=${tmpline%%_*} tmpline=${tmpline#*_}
    d2=${tmpline%%_*} tmpline=${tmpline#*_}
    d3=${tmpline%%_*} tmpline=${tmpline#*_}
    d4=${tmpline%%@*} tmpline=${tmpline#*@}
    d5=${tmpline%%,*} tmpline=${tmpline#*,}
    d6=${tmpline}

    printf "%s\n" "d1=$d1" "d2=$d2" "d3=$d3" "d4=$d4" "d5=$d5" "d6=$d6"
done 

In bash, I would use a regular expression instead.

$ cat input
one_two_three_fourpt1_fourpt2@fivept1_fivept2,six
$ regex='([^_]+)_([^_]+)_([^_]+)_(.+)@([^,]+).(.*)'
$ while IFS= read -r line; do
> [[ $line =~ $regex ]]
> done < input
$ printf '%s\n' "${BASH_REMATCH[@]}"
one_two_three_fourpt1_fourpt2@fivept1_fivept2,six
one
two
three
fourpt1_fourpt2
fivept1_fivept2
six

Element zero of BASH_REMATCH contains the entire match; the remaining elements contain the individual capture groups, from the left.

Alternatively, you can use read to split first on @, then again to split the two halves using _ and , as appropriate.

$ IFS="@" read -r first second <<< "$line"
$ IFS=_ read -r f1 f2 f3 f4 <<< "$first"
$ IFS=, read -r f5 f6 <<< "$second"

Because the second call to read only has 4 arguments, f4 will contain whatever follows the 3rd _, without any further field splitting on additional _s.


A similar regular expression and two-level splitting scheme can be used in a language that supports more efficient iteration over the contents of a file, which (as Nahuel Fouilleul points out) bash doesn't do very quickly. (read reads its input byte-by-byte, rather than reading entire chunks at once, to avoid reading more bytes than necessary to consume exactly one line of input.)

If you have more than 1 time a field with @._._....,
You can try this awk :

echo "data1_data2@d_a_t_a_17,data3_data4@d_a_t_a_5,data6_data7" |
awk '
{
i = split ( $0 , a , "_" )
for ( j = 1 ; j <= i ; j++ )
  if ( a[j] !~ /@/ )
    print "d" ++k "==\"" a[j] "\""
  else
    {
      split ( a[j] , b , "@" )
      print "d" ++k "==\"" b[1] "\""
      sub ( ".*@" , "" , a[j] )
      while ( a[j] !~ "," )
        {
          c = c a[j] "_"
          j++
        }
        split ( a[j] , b , "," )
        c = c b[1]
        print "d" ++k "==\"" c "\""
        a[j] = b[2]
        j--
        c = ""
    }
}'
Related