Input:
#CHROM POallelic REF ALT QUAL FILTER INFO FORMAT SAMPLE1
10 111 . A ACCA,TCGG . PASS AC=4,2;AF=0.345,0.25;AN=8 GT:AD:DP:GQ:PL 0/1:2,15,31:30:99:2407,0,533,697,822,574
Expected output:
#CHROM POS ID REF ALT QUAL FILTER INFO FORMAT SAMPLE1
10 111 . A ACCA . PASS AC=4;AF=0.345;AN=8;OLD_MULTIALLELIC=20:101:A/ACCA/TCGG GT:AD:DP:GQ:PL 0/1:2,15:30:99:2407,0,533
10 111 . A TCGG . PASS AC=2;AF=0.25;AN=8;OLD_MULTIALLELIC=20:101:A/ACCA/TCGG GT:AD:DP:GQ:PL 0/.:2,31:30:99:2407,697,574
Code: import os import sys
def main():
PATH_Annotation_file='./XXXX.vep.vcf'
PATH_output='./XXX.txt'
parser_annotation_file = ParseAnnotationFile(PATH_Annotation_file, PATH_output)
file = parser_annotation_file.extract_allele_information()
print(file)
class ParseAnnotationFile:
def __init__(self,
path_Annotation_file,
path_output):
self.path_Annotation_file = path_Annotation_file
self.path_output = path_output
def extract_allele_information(self):
path_Annotation_file = self.path_Annotation_file
bool_start_variant = False
with open(path_Annotation_file, "r") as vep_vcf:
for line in vep_vcf:
if "#CHROM" in line:
bool_start_variant = True
continue
elif bool_start_variant:
list_variant = line.split('\t')
info = list_variant[7]
list_info = info.split(';')
print(list_info[0][3:]) # Allele_Count
print(list_info[1][3:]) # Allele_Frequency
print(list_info[2][3:]) # Allele_Number
if __name__ == "__main__":
main()
I'm trying to split multiallelic row to biallelic rows in variant annotation file(.vep.vcf). However, I don't have ideas to duplicate rows and write other materials in 'ALT' column, 'INFO column's AC,AF,AN' each row. Please let me know it if you have a good idea. Thank you!