I am following this guide to get some data from a protein file using the ESM-1b Transformer.
I have created a list of tuples out of a Fasta file that has some protein IDs and their sequences in each tuple.
parsed_list = [(k, v) for k, v in parsed_seqs.items()]
However, I ran out of RAM so I sliced the list in 650 samples:
new_list = parsed_list[:650]
When following the Guide I encounter the following error `TypeError: unhashable type: 'slice' when trying to append the model tokens to another list. How should I be accessing the sliced list?
The code I have for now:
batch_labels, batch_strs, batch_tokens = batch_converter(new_list)
with torch.no_grad():
results = model(batch_tokens, return_contacts=True)
token_representations = results["representations"]
sequence_representations = []
for i, (_, seq) in enumerate(new_list): #Here is where I get the error
sequence_representations.append(token_representations[i, 1 : len(seq) + 1].mean(0))