For an NLP task, given a mapping, I need to encode each unicode character to an integer in a list of list of words. I'm trying to figure out a quick way to do this without dropping down into cython.
Here's a slowish way to write the function:
def encode(sentences, mappings):
encoded_sentences = []
for sentence in sentences:
encoded_sentence = []
for word in sentence:
encoded_word = []
for ch in word:
encoded_word.append(mappings[ch])
encoded_sentence.append(encoded_word)
encoded_sentences.append(encoded_sentence)
return encoded_sentences
Given the following input data:
my_sentences = [['i', 'need', 'to'],
['tend', 'to', 'tin']]
mappings = {'i': 0, 'n': 1, 'e': 2, 'd':3, 't':4, 'o':5}
I want encode(my_sentences, mappings) to produce:
[[[0], [1, 2, 2, 3], [4, 5]],
[[4, 2, 1, 3], [4, 5], [4, 0, 1]]]