I've got a file with a subset of geneIDs in it, and a fasta file with all geneIDs and their sequences. For each gene in the subset file, I want to get positions 2-7 from the start of each fasta sequence. Ideally the output file would be 'pos 2-7' '\t' 'geneID'.
Example subset:
mmu-let-7g-5p MIMAT0000121
mmu-let-7i-5p MIMAT0000122
Fasta file:
>mmu-let-7g-5p MIMAT0000121
UGAGGUAGUAGUUUGUACAGUU
>mmu-let-7i-5p MIMAT0000122
UGAGGUAGUAGUUUGUGCUGUU
>mmu-let-7f-5p MIMAT0000525
UGAGGUAGUAGAUUGUAUAGUU
wanted output:
GAGGUA mmu-let-7g-5p MIMAT0000121
GAGGUA mmu-let-7i-5p MIMAT0000122
The first part (pulling out fasta sequences for the subset of genes) I've done using grep -w -A 1 -f. Not sure how to get pos 2-7 and make the output look like that now using Bash.