I've spent the day trying to figure this out but didn't succeed. I have two files like this:
File1:
chr id pos
14 ABC-00 123
13 AFC-00 345
5 AFG-99 988
File2:
index id chr
1 ABC-00 14
2 AFC-00 11
3 AFG-99 7
I wanna check if the value for chr in File 1 != from chr in File 2 for the same ID, if this is true I want to print some columns from both files to have an output like this one below.
Expected output file:
ID OLD_chr(File1) NEW_chr(File2)
AFC-00 13 11
AFG-99 5 7
.....
Total number of position changes: 2
I got one caveat though. In File 1 I have to substitute some values in the $1 column before comparing the files. Like this:
30 and 32 >> X
31 >> Y
33 >> MT
Because in File 2 that's how those values are coded. And then compare the two files. How in the hell can I achieve this?
I've tried to recode File 1:
awk '{
if($1=30 || $1=32) gsub(/30|32/,"X",$1);
if($1=31) gsub(/31/,"Y",$1);
if($1=33) gsub(/33/,"MT",$1);
print $0
}' File 1 > File 1 Recoded
And I was trying to match the columns and print the output with:
awk 'NR==FNR{a[$1]=$1;next} (a[$1] !=$3){print $2, a[$1], $3 }' File 1 File 2 > output file