How to check delimiter of txt file in ubuntu

Viewed 704

I have .txt file and I want to check their delimiter. File may have delimiter tab, pipe(|) and comma

Tab separated file

ID  Name    Email
1   Test    a@test.com
2   testone b@test.com

Comma separated

ID,Name,Email
1,Test,a@test.com
2,testone,b@test.com

For above sample data I want to get delimiter. So for first sample delimiter is tab and for second delimiter is comma(,)

3 Answers

This is actually a very good question and I'd like to see other peoples' solutions as well. This is something one needs to wrestle with when automizing processing of sh***y data:

awk '
FNR==1 {                          # process the header record
    line=$0                       # duplicate to leave $0 usable
    
    gsub(/[^,|\t]/,"",line)       # remove non-candidates

    split(line,a,"")              # split leftovers

    delete b                      # ... since FNR...
    max=prev=0                    # reset
        
    for(i in a)                   # flip a and count hits
        b[a[i]]++
        
    for(i in b)                   # find max amount of hits
        if(b[i]>=b[max]) {        
            prev=max
            max=i
        }
    if(b[prev]==b[max]) {         # if count collision
        print "Multiple candidates for delimiter. Exiting."
        exit 1
    }
                                  # below: output 
    printf "Delimiter: %s\n",(max=="\t"?"\\t":(max==" "?"[space]":max))

    exit
}' file

Output for example:

Delimiter: \t

You cannot be sure, consider this:

test,test|test,test|test
test,test|test,test|test

But you can have a good estimation. I would parse a sample of the file, e.g. first 10 lines with a script like this:

BEGIN {
   sep[","]   = "comma"
   sep["\\|"] = "pipe"
   sep["\t"]  = "tab"
}

{
    for (x in sep) {
        c = gsub(x,"&",$0)
        if (c) cnt[sep[x] " " (c+1)]++
    }
}

END {
    for (x in cnt) {
        if (max == "" || cnt[x] > max) {
            max = cnt[x]
            est = x
        }
    }
    print est
}

where we count lines where a candidate delimiter appeared for N times, excluding zero times, and we keep as the best guess whatever happened most in the file.

Example usage:

head file | awk -f tst.awk

It prints the estimated separator and the estimated number of fields. Example:

> cat file
1,2,3
1,2,3
> awk -f tst.awk file
comma 3

The rule how to find the delimiter is unclear. The first delimiter found is given by

# Substitute a tab for the space in the next command
grep -Eom1 "[,| ]" file | head -1

Maybe you want the character after ID:

sed -rn '1s/..(.).*/\1/p' file
# Or after a word
sed -rn '1s/\w*(.).*/\1/p' file
Related