I've been task with kind of summarising a few files into a tsv file. I have to select specific data from a list of files and write it as a line of tab-seperated columns in a tsv file. Every line in the files have a 'name' as a first column so it is easy to filter data ($1 == "NAME"). One file == one line in tsv. So far I wrote this:
#! /bin/bash
cat > newFile.txt
for f in *.pdb; do
awk '$1 == "ACCESSION" {print $2}' ORS="/t" "$f" >> newFile.txt
awk '$1 == "DEFINITION" {print $2}' ORS="/t" "$f" >> newFile.txt
awk '$1 == "SOURCE" {print $2}' ORS="/t" "$f" >> newFile.txt
awk '$1 == "LOCUS" {print$4}' ORS="/r" "$f" >> newFile.txt
done
Obviously this attrocity of a code does not work. Is it possible to modify what I wrote and complete the task using awk?
Example of a file:
LOCUS \t NM_123456 \t 2000bp \t mRNA
DEFINITION \t Very nice gene from a very nice mouse
ACCESSION \t NM_123456
VERSION \t 1.000
SOURCE \t Very nice mouse
end result:
NM_123456 /t Very nice gene from a very nice mouse /t Very nice mouse /t mRNA
NM_345678 /t Not so nice gene from an angry elephant /t Angry Elephant /t mRNA
"/t" stands for a tab (I did not know how to write it down sorry). Also the example files contain much more information, I just gave a 'header' let's say.