String split and extract the last field in bash

Viewed 157

I have a text file FILENAME. I want to split the string at - of the first column field and extract the last element from each line. Here "$(echo $line | cut -d, -f1 | cut -d- -f4)"; alone is not giving me the right result.

FILENAME:

TWEH-201902_Pau_EX_21-1195060301,15cef8a046fe449081d6fa061b5b45cb.final.cram
TWEH-201902_Pau_EX_22-1195060302,25037f17ba7143c78e4c5a475ee98e25.final.cram
TWEH-201902_Pau_T-1383-1195060311,267364a6767240afab2b646deec17a34.final.cram

code I tried:

while read line; do \
DNA="$(echo $line | cut -d, -f1 | cut -d- -f4)";
echo $DNA
done < ${FILENAME} 

Result I want

1195060301
1195060302
1195060311
5 Answers

Would you please try the following:

while IFS=, read -r f1 _; do    # set field separator to ",", assigns f1 to the 1st field and _ to the rest
    dna=${f1##*-}               # removes everything before the rightmost "-" from "$f1"
    echo "$dna"
done < "$FILENAME"

I do not know the constraints on your input file, but if what you are looking for is a 10-digit number, and there is only ever one 10-digit number per line... This should do niceley

grep -Eo '[0-9]{10,}' input.txt
1195060301
1195060302
1195060311

This essentially says: Show me all 10 digit numbers in this file

input.txt

TWEH-201902_Pau_EX_21-1195060301,15cef8a046fe449081d6fa061b5b45cb.final.cram
TWEH-201902_Pau_EX_22-1195060302,25037f17ba7143c78e4c5a475ee98e25.final.cram
TWEH-201902_Pau_T-1383-1195060311,267364a6767240afab2b646deec17a34.final.cram

Well, I had to do with the two lines of codes. May be someone has a better approach.

while read line; do \
DNA="$(echo $line| cut -d, -f1| rev)"
DNA="$(echo $DNA| cut -d- -f1 | rev)"
echo $DNA
done < ${FILENAME}

A sed approach:

sed -nE 's/.*-([[:digit:]]+)\,.*/\1/p' input_file

sed options:

  • -n: Do not print the whole file back, but only explicit /p.
  • -E: Use Extend Regex without need to escape its grammar.

sed Extended REgex:

  • 's/.*-([[:digit:]]+)\,.*/\1/p': Search, capture one or more digit in group 1, preceded by anything and a dash, followed by a comma and anything, and print only the captured group.

Using awk:

awk -F[,] '{ split($1,arr,"-");print arr[length(arr)] }' FILENAME

Using , as a separator, take the first delimited "piece" of data and further split it into an arr using - as the delimiter and awk's split function. We then print the last index of arr.

Related