I have a file of peptide sequences in peptides.txt that I would like to match to my protein database human_proteins.fasta. I would like to match the peptide list to the protein database and fetch the protein ID which is in the preceding line. Some of the peptides have multiple matches to the protein database.
Ultimately, I would like to produce a table/dataframe like this:
| Peptide | No. of matches | Sequence ID | Protein Sequence |
|---|---|---|---|
| AAAAA | 2 | ENST0001 | AAAAABCFMED |
| AAAAA | 2 | ENST0002 | AAAAAXXX |
The first few lines of my hypothetical protein database human_proteins.fasta look like this:
>ENST0001
AAAAABCFMED
>ENST0002
AAAAAXXX
>ENST0003
MGRVSGLVPSR
peptides.txt looks like this:
AAAAA
LSSPATLNSR
HETLTSLNLEK
GGGGNFGPGPGSNFR
VSEQGLIEILK
DFLAGGIAAAISK
I am using the following command in bash
while read line; do printf $line grep -B 1 $line ../databases/human_proteins.fasta < peptides.txt
and I am able to get output like this:
>ENST0001
AAAAABCFMED
>ENST0002
AAAAAXXX
However, I am having trouble processing the output into a table. Is there a nice solution in unix/bash that can solve this?