awk count specified char occur , return 0 if not found

Viewed 86

I have a textfile:

and

b

,

.

apple

banana

and I want to count the occurrences of some specific characters which include semi_colon, in this case there's no semi_colon found

excepted output would be : semi_colon 0

Here's my code:

sed -e 's/;/semi_colon/g' data.txt|awk '{count[$1]++} if(count[$1]==""){count[$1]=0} END{print"semi_colon",count["semi_colon"]}'

which gets the output like :

semi_colon

Wondering how to achieve the expected output:

semi_colon 0

Cheers if anyone can help!

7 Answers

With your shown samples, please try following awk code.

awk '
{
  countA+=gsub(/a/,"&")
  countSemiColon+=gsub(/;/,"&")
}
END{
  print "a "countA+0 ORS "semi_colon " countSemiColon+0
}
'  Input_file

Explanation: Simple explanation would be, in main block of awk program using awk's gsub function/method to globally substitute a with itself(just for counting its occurrences sake) and putting its number of substitution occurrences into countA awk variable, where += denotes that its value will keep adding to its previous value itself(to get TOTAL number of a's in all lines).

Then creating variable named countSemiColon which has value of global substitutions of ; in each line and it keep adding its all values from all lines to get all occurrences in whole Input_file. In END block of awk printing the value of variables countA and countSemiColon as per required output.

Using all awk to count characters in file:

$ awk 'BEGIN {
    c[";"]="semi_colon"    # define chars to count and their nicknames
    c["a"]="a"
}
{
    split($0,t,"")         # split record between chars
    for(i in t)            # iterate every char in string
        f[t[i]]++          # count freqs
}
END {
    for(i in c)
        print c[i],f[i]+0  # +0 is the key to fix space to a zero
}' file

Output:

a 5
semi_colon 0
$ cat tst.awk
{
    for ( i=1; i<=length($0); i++ ) {
        cnt[substr($0,i,1)]++
    }
}
END {
    print "semi_colon", cnt[";"]+0
}

$ awk -f tst.awk file
semi_colon 0

and for multiple chars:

$ cat tst.awk
{
    for ( i=1; i<=length($0); i++ ) {
        cnt[substr($0,i,1)]++
    }
}
END {
    chars = "a;."
    map[";"] = "semi_colon"
    for ( i=1; i<=length(chars); i++ ) {
        char = substr(chars,i,1)
        print (char in map ? map[char] : char), cnt[char]+0
    }
}

$ awk -f tst.awk file
a 5
semi_colon 0
. 1

Regarding your original code sed -e 's/;/semi_colon/g' data.txt|awk... - you never need sed when you're using awk and doing that would break your script if you also wanted it to count any of the letters that are present in the string semi_colon (e.g. _s or es) or if the input contained a string that was semi_colon as you'd then have no way to differentiate that from an original ;. In general, when desirable you should map symbols to strings in the output, not in the input.

First, your awk script is not correct (typo?)

awk '
    {
        count[$1]++
    } 
    if(count[$1]=="")
    {
        count[$1]=0
    }
    END{
        print"semi_colon",count["semi_colon"]
    }
'

It should be

awk '
    {
        count[$1]++
        if(count[$1]=="")
        {
            count[$1]=0
        }
    }
    END{
        print"semi_colon",count["semi_colon"]
    }
'

That said, the

if(count[$1]=="")
{
    count[$1]=0
}

Is never called: since $1 must have been matched when arriving here, count[b1] can't be "" (it's default value).

The idea is right, but you need to insert it in END block:


awk '
    {
        count[$1]++
    }
    END{
        if (count["semi_colon"] != "") {
            print"semi_colon",count["semi_colon"]
        } else 
        {
            print"semi_colon", 0
        }
    }
'
echo -e "and\nb\n,\napple\nbanana" | sed -e 's/;/semi_colon/g' |awk '  { count[$1]++  } END{ if (count["semi_colon"] != "") {  print"semi_colon",count["semi_colon"]  } else   {  print"semi_colon", 0  }  } '

But use the other answer

Firstly you have typo preventing your code from execution

awk '{count[$1]++} if(count[$1]==""){count[$1]=0} END{print"semi_colon",count["semi_colon"]}'

should be

awk '{count[$1]++}count[$1]==""{count[$1]=0} END{print"semi_colon",count["semi_colon"]}'

more importantly count[$1]=="" never hold true, as it was done after count[$1]++ meaning that during that check count[$1] is always integer equal or greater than 1. Therefore above code will always gives same output as

awk '{count[$1]++}END{print"semi_colon",count["semi_colon"]}'

As you are dealing with numeric values you might use +0 as this used against empty string gives 0 and is neutral against known numeric values. If you wish more general way to provide placeholder for unknown value then combine in check with so called ternary operator, consider following example let file.txt content be

A
A
A

and you want to count A, B, C and output not found for these which did not exist, then you might use GNU AWK following way

awk '{cnt[$1]+=1}END{print "A:",("A" in cnt)?cnt["A"]:"not found","B:",("B" in cnt)?cnt["B"]:"not found","C:",("C" in cnt)?cnt["C"]:"not found"}' file.txt

output

A: 3 B: not found C: not found

Explanation: in is used for checking if array has key, so-called ternary operator has form of condition?valueiftrue:valueiffalse

(tested in gawk 4.2.1)

{m/n/g}awk '
BEGIN {    ___ = " char::{ %3s } > freq::{ %5.f }\n"
           FS  = substr(_, OFS = _,   _ = ";")"a"
}  NF { __[_] +=   gsub(_,"&")+(__[FS] += --NF)*_
} END { 
       for(_ in __) { printf(___,_,__[_]) } }'

————————————————————————————————————————

 char::{   ; } > freq::{     0 }
 char::{   a } > freq::{     5 }

if you want a full-fledged character freq counter, I just quickly slapped one together for UTF-8 :::

  • it'll still work with other awk variants, just without the unicode bit

  • even gawk in unicode mode can be used to count arbitrary bytes in binary files without triggering any error messages

    • I even threw an mp3 file at it

-- unfortunately, a small bug remains - it's mis-classifying the UTF-16 surrogate range

  • in terms of performance, if you use mawk 1.9.9.6/mawk2, I just freq-counted a 408 MB text file filled with Unicode in 21.7secs :

————————————————

 char >>[    ? ]--[  \222 \x92 |     146 (8-bit byte)   ] freq >>[      128404 ] sumtotal >>[      405114168 ]
 char >>[    7 ]--[  U+     37 |      55 (ASCII)        ] freq >>[     4197205 ] sumtotal >>[      409311373 ]
 char >>[    H ]--[  U+     48 |      72 (ASCII)        ] freq >>[      212959 ] sumtotal >>[      409524332 ]
 char >>[    ? ]--[  \355 \xED |     237 (8-bit byte)   ] freq >>[     4909733 ] sumtotal >>[      414434065 ]
 char >>[    ? ]--[  \233 \x9B |     155 (8-bit byte)   ] freq >>[      747521 ] sumtotal >>[      415181586 ]
 char >>[    < ]--[  U+     3C |      60 (ASCII)        ] freq >>[       42849 ] sumtotal >>[      415224435 ]
 char >>[    Q ]--[  U+     51 |      81 (ASCII)        ] freq >>[        8615 ] sumtotal >>[      415233050 ]
 char >>[    ? ]--[  \352 \xEA |     234 (8-bit byte)   ] freq >>[     8878221 ] sumtotal >>[      424111271 ]
 char >>[    ? ]--[  \240 \xA0 |     160 (8-bit byte)   ] freq >>[     4432353 ] sumtotal >>[      428543624 ]
( pvE 0.1 in0 < "${m3r}" | mawk2 ; )  21.59s user 0.25s system 100% cpu 21.734 total

    gawk/mawk 'BEGIN {
     1      _ = initEscRE(____) * split(FS = _, __, _)
     1      FS = "^$"
    }

    # Rule(s)

 50355  {
5049452     while (/./) {
5049452         ___ = sprintf("%.*s", ! _, $_)
5049452         __[___] += gsub(___ in ____ ? ____[___] : ___, "", $0)
        }
    }

    # END rule(s)

    END {
     1      initCHR(__, _____)
     1      PROCINFO["sorted_in"] = "@ind_str_asc"
     1      ___ = ""
     1      OFMT = "%.f"
 18964      for (_ in __) {
 18964          print sprintf(" char >>[ %4s ]--[ %-35s ] freq >>[ %11.f ] sumtotal >>[ %14.f ]", _, _____[_], __[_], ___ += __[_])
        }
    }


    # Functions, listed alphabetically

     1  function initCHR(__, ___, _, ____, _____, ______)
    {
     1      split(_, ___, FS)
     1      ______ *= ______ = ____ = (_____ = _ = 4) ^ _
     1      _____ *= _____
1114112     for (_ = ______ - (-_ < +_) + ______ * _____; -_ <= +_; _--) {
1114112         ___[sprintf("%c", _)] = sprintf(" U+ %6X | %7.f (%s)", _, _, (_ + _) < (_____ * _____) ? "ASCII" : int((_ + _) / (____ * _____)) ~ "^5[45]$" ? "UTF16 surrogates" : "UTF8 " (4 - (_ < ______) - ((_ + _) < (____ * _____))) "-bytes")
        }
     1      ____ ^= ____ = (______ = _ *= _ += _ ^= _ < _) + _
     1      _ += _ = _ * _ * -_
     1      ______ ^= ______
   128      while (_ < -_) {
   128          ___[sprintf("%c", _ + ____)] = sprintf(" \\%o \\x%2X | %7.f (8-bit byte)", _ + ______, _ + ______, ______ + _++)
        }
    }

     1  function initEscRE(__, ___, _, ____)
    {
     1      split(_, __, FS) + split("!\"#$%&'()*+,-./:;<=>?@[\\]^_`{|}~", ___, _)
    32      for (____ in ___) {
    32          _ = ___[____]
    32          __[_] = ("[") (index("^]\\/", _) ? "\\" : "") (_) ("]")
        }
     1      return (_ < _)
    }

———————————————————————————

 char >>[    s ]--[  U+     73 |     115 (ASCII)        ] freq >>[       29805 ] sumtotal >>[        4706098 ]
 char >>[    t ]--[  U+     74 |     116 (ASCII)        ] freq >>[       32722 ] sumtotal >>[        4738820 ]
 char >>[    u ]--[  U+     75 |     117 (ASCII)        ] freq >>[       27833 ] sumtotal >>[        4766653 ]
 char >>[    v ]--[  U+     76 |     118 (ASCII)        ] freq >>[       27453 ] sumtotal >>[        4794106 ]
 char >>[    w ]--[  U+     77 |     119 (ASCII)        ] freq >>[       21278 ] sumtotal >>[        4815384 ]
 char >>[    x ]--[  U+     78 |     120 (ASCII)        ] freq >>[       31107 ] sumtotal >>[        4846491 ]
 char >>[    y ]--[  U+     79 |     121 (ASCII)        ] freq >>[       28413 ] sumtotal >>[        4874904 ]
 char >>[    z ]--[  U+     7A |     122 (ASCII)        ] freq >>[       27555 ] sumtotal >>[        4902459 ]
 char >>[    { ]--[  U+     7B |     123 (ASCII)        ] freq >>[       20195 ] sumtotal >>[        4922654 ]
 char >>[    | ]--[  U+     7C |     124 (ASCII)        ] freq >>[       24009 ] sumtotal >>[        4946663 ]
 char >>[    } ]--[  U+     7D |     125 (ASCII)        ] freq >>[       21986 ] sumtotal >>[        4968649 ]
 char >>[    ~ ]--[  U+     7E |     126 (ASCII)        ] freq >>[       22780 ] sumtotal >>[        4991429 ]
 char >>[      ]--[  U+     7F |     127 (ASCII)        ] freq >>[       21827 ] sumtotal >>[        5013256 ]
 char >>[    ? ]--[  \200 \x80 |     128 (8-bit byte)   ] freq >>[       42532 ] sumtotal >>[        5110931 ]
 char >>[    ? ]--[  \201 \x81 |     129 (8-bit byte)   ] freq >>[       44412 ] sumtotal >>[        5155343 ]
 char >>[    ? ]--[  \202 \x82 |     130 (8-bit byte)   ] freq >>[       36002 ] sumtotal >>[        5191345 ]
 char >>[    ? ]--[  \203 \x83 |     131 (8-bit byte)   ] freq >>[       42798 ] sumtotal >>[        5234143 ]
 char >>[    ? ]--[  \204 \x84 |     132 (8-bit byte)   ] freq >>[       37037 ] sumtotal >>[        5271180 ]
 char >>[    ? ]--[  \205 \x85 |     133 (8-bit byte)   ] freq >>[       41117 ] sumtotal >>[        5312297 ]
 char >>[    ? ]--[  \206 \x86 |     134 (8-bit byte)   ] freq >>[       31655 ] sumtotal >>[        5343952 ]
 char >>[    ? ]--[  \207 \x87 |     135 (8-bit byte)   ] freq >>[       43455 ] sumtotal >>[        5387407 ]
 char >>[    ? ]--[  \210 \x88 |     136 (8-bit byte)   ] freq >>[       37021 ] sumtotal >>[        5424428 ]
 char >>[    ? ]--[  \211 \x89 |     137 (8-bit byte)   ] freq >>[       34333 ] sumtotal >>[        5458761 ]

 char >>[    ɣ ]--[  U+    263 |     611 (UTF8 2-bytes) ] freq >>[         112 ] sumtotal >>[        7412934 ]
 char >>[    ɤ ]--[  U+    264 |     612 (UTF8 2-bytes) ] freq >>[         118 ] sumtotal >>[        7413052 ]
 char >>[    ɥ ]--[  U+    265 |     613 (UTF8 2-bytes) ] freq >>[         104 ] sumtotal >>[        7413156 ]
 char >>[    ɦ ]--[  U+    266 |     614 (UTF8 2-bytes) ] freq >>[         101 ] sumtotal >>[        7413257 ]

 char >>[    ࠁ ]--[  U+    801 |    2049 (UTF8 3-bytes) ] freq >>[           1 ] sumtotal >>[        8089676 ]
 char >>[    ࠂ ]--[  U+    802 |    2050 (UTF8 3-bytes) ] freq >>[           3 ] sumtotal >>[        8089679 ]
 char >>[    ࠈ ]--[  U+    808 |    2056 (UTF8 3-bytes) ] freq >>[           2 ] sumtotal >>[        8089681 ]
 char >>[    ࠉ ]--[  U+    809 |    2057 (UTF8 3-bytes) ] freq >>[           1 ] sumtotal >>[        8089682 ]
 char >>[    ࠌ ]--[  U+    80C |    2060 (UTF8 3-bytes) ] freq >>[           1 ] sumtotal >>[        8089683 ]

 char >>[     ]--[  U+  27569 |  161129 (UTF8 4-bytes) ] freq >>[           1 ] sumtotal >>[        8539054 ]
 char >>[     ]--[  U+  28041 |  163905 (UTF8 4-bytes) ] freq >>[           1 ] sumtotal >>[        8539055 ]
 char >>[     ]--[  U+  282A8 |  164520 (UTF8 4-bytes) ] freq >>[           1 ] sumtotal >>[        8539056 ]
 char >>[     ]--[  U+  286F1 |  165617 (UTF8 4-bytes) ] freq >>[           1 ] sumtotal >>[        8539057 ]
 char >>[     ]--[  U+  2882A |  165930 (UTF8 4-bytes) ] freq >>[           1 ] sumtotal >>[        8539058 ]
Related