I have three very large files and I just want to grab the matching ID2s that come up for the first 20 ID2s for file 1 (in the example, I listed only 5). I want it to search file 2 and file 3 in its entirety and pull out the lines that match to the 20 ID2s from file1. I also want the ID1- even if it's blank (as well as other columns that are on the same matching ID2 line).
I know how to do this in python, but since the files are so large, my computer can't handle it. I have been trying to do it in the terminal (unix) with no luck. I tried the following command but 1)it doesn't deal with three files and 2) it is also throwing me an error (maybe I'm grabbing the wrong columns)?
awk ‘NR==FNR{a[$1][$0];next} $0 in a {for (i in a[$0]) print i}’ file1.tsv file2.tsv
Error:
zsh: parse error near `}'
Another question: How do I skip the first five rows in unix when doing the command (because these files have some notes in the beginning that are read in as rows)?
file1
id 1 id2 name
A00000004 B00000004 emily
B00000005 joe
A00000006 B00000006 jack
B00000007 john
A00000008 B00000008 sally
file2
id1 id2 source source_code
A00000001 B00000001 source1 321
A00000003 B00000003 source2 165
A00000004 B00000004 source3 481
A00000005 B00000005 source2 891
file 3
id2 code
B00000006 yes
B00000007 no
B00000005 yes
B00000004 yes
B00000012 yes
Desired:
id1 id2 name source source_code code
A00000004 B00000004 emily source3 481 yes
A00000005 B00000005 joe source2 891 yes