I am using PuTTy for school to learn UNIX/Linux and have a file 2.asr which is a large data set containing the age, sex, and race of multiple individuals in their own columns, for example:
19 Male White
23 Female White
23 Male White
45 Female Other
54 Male Asian
24 Male Other
34 Female Asian
23 Male Hispanic
45 Female Hispanic
38 Female White
I would like to find the average age, max age, min age, and total occurrences of unique demographics such as Male White or Female Hispanic.
I've tried using awk code as follows:
$ awk '$2 == "Male" && $3 == "Hispanic" {sum+=$1; n++}
(NR==1) {min=$1;max=$1+0};
(NR>=2) {if(min>$1) min=$1; if(max<$1) max=$1}
END {if (n>0)
print $2 " " $3 " Average Age: " sum/n ", Max: " max ", Min: " min ", Total: " n
}' 2.asr
However, regardless of what sex and race I input, the output is always "Male White" and the max and min values are those of the entire dataset rather than the unique demographic conditions I've set. It does seem however that the average age and total occurrences of each demographic are outputted properly and change accordingly. I've tried using $2 and $3 at the start of the command in an if statement and utilizing BEGIN at the start also but I keep getting syntax errors at the end where I have my print function. Is there a better way to approach this with if statements ate the start of the command or is my syntax off somewhere? Thanks to whoever wishes to assist!