awk print sum of group of lines

Viewed 103

I have a file with a column named (effect) which has rows separated by blank lines,

(effect)
    1
    1
    1
    
    (effect)
    1
    1
    1
    1
    
    
    (effect)
    1
    1

I know how to print the sum of column like

awk  '{sum+=$1;} END{print sum;}' file.txt

Using awk how can I print the sum of each (effect) in for loop? such that I have three lines or multiple lines in other cases like below

sum=3
 sum=4
 sum=2
5 Answers

With your shown samples, please try following awk code. Written and tested in GNU awk.

awk -v RS='(^|\n)?\\(effect\\)[^(]*' '
RT{
  gsub(/\(effect\)\n|\n+[[:space:]]*$/,"",RT)
  num=split(RT,arr,ORS)
  print "sum="num
}
'  Input_file

Explanation: Simple explanation would be, using GNU awk. In awk program set RS as (^|\n)?\\(effect\\)[^(]* regex for whole Input_file. In main program checking condition if RT is NOT NULL then using gsub(Global substitution) function to substitute (effect)\n and \n+[[:space:]]*$(new lines followed by spaces at end of value) with NULL in RT. Then splitting value of RT into array named arr with delimiter of ORS and saving its(total contents value OR array length value) into variable named num, then printing sum= along with value of num here to get required results.

With shown samples, output will be as follows:

sum=3
sum=4
sum=2

You can check if there is an (effect) part, and print the sum when encountering either the (effect) part or when in the END block.

awk '
$1 == "(effect)" { if(seen) print "sum="sum; seen = 1; sum = 0 }
/[0-9]/ { sum += $1 }
END { if (seen) print "sum="sum }
' file

Output

sum=3
sum=4
sum=2

This should work in any version of awk:

awk '{sum += $1} $0=="(effect)" && NR>1 {print "sum=" sum; sum=0} 
END{print "sum=" sum}' file

sum=3
sum=4
sum=2

Similar to @Ravinder's answer, but does not depend on the name of the header:

awk -v RS='' -v FS='\n' '{
    sum = 0
    for (i=2; i<=NF; i++) sum += $i
    printf "sum=%d\n", sum
}' file

RS='' means that sequences of 2 or more newlines separate records.
The Field Separator is newline.
The for loop omits field #1, the header.

However that means that empty lines truly need to be empty: no spaces or tabs allowed. If your data might have blank lines that contain whitespace, you can set

-v RS='\n[[:space:]]*\n'
$ awk -v RS='(effect)' 'NR>1{sum=0; for(i=1;i<=NF;i++) sum+=$i; print "sum="sum}' file

sum=3
sum=4
sum=2
Related