Fastest way to encode characters in a list of list of strings

Viewed 492

For an NLP task, given a mapping, I need to encode each unicode character to an integer in a list of list of words. I'm trying to figure out a quick way to do this without dropping down into cython.

Here's a slowish way to write the function:

def encode(sentences, mappings):
    encoded_sentences = []
    for sentence in sentences:
        encoded_sentence = []
        for word in sentence:
            encoded_word = []
            for ch in word:
                encoded_word.append(mappings[ch])
            encoded_sentence.append(encoded_word)
        encoded_sentences.append(encoded_sentence)
    return encoded_sentences

Given the following input data:

my_sentences = [['i', 'need', 'to'],
            ['tend', 'to', 'tin']]

mappings = {'i': 0, 'n': 1, 'e': 2, 'd':3, 't':4, 'o':5}

I want encode(my_sentences, mappings) to produce:

[[[0], [1, 2, 2, 3], [4, 5]],
 [[4, 2, 1, 3], [4, 5], [4, 0, 1]]]
2 Answers
Related