I'm trying to transfer a .fasta file into a .xls file so that I can conveniently color my phylogenetic tree.
import pandas as pd
import re
from Bio import SeqIO
s1 = {}
s2 = {}
with open('/Users/xxx.fasta') as seqF:
for seqFP in SeqIO.parse(seqF,"fasta"):
seq_id = seqFP.description
seq_an = re.search(r"\[([^]]*)\]",seq_id)
s1[seqFP.id] = seqFP.seq
s2[seqFP.id] = seq_an
print(s2.values)
However when I tried to use the s1 to create a pd.Series, the sequence column shows like "(M, S, A, C, C, N, K, L, A, V, L, G, L, T, F, ...".