I am loading a TextLineDataset and I want to apply a tokenizer trained on a file:
import tensorflow as tf
data = tf.data.TextLineDataset(filename)
MAX_WORDS = 20000
tokenizer = Tokenizer(num_words=MAX_WORDS)
tokenizer.fit_on_texts([x.numpy().decode('utf-8') for x in train_data])
Now I want to apply this tokenizer on data so that each word is replaced with its encoded value. I have tried data.map(lambda x: tokenizer.texts_to_sequences(x)) which gives OperatorNotAllowedInGraphError: iterating over tf.Tensor is not allowed in Graph execution. Use Eager execution or decorate this function with @tf.function.
Following the instruction, when I write the code as:
@tf.function
def fun(x):
return tokenizer.texts_to_sequences(x)
train_data.map(lambda x: fun(x))
I get: OperatorNotAllowedInGraphError: iterating over tf.Tensor is not allowed: AutoGraph did convert this function. This might indicate you are trying to use an unsupported feature.
So how to do the tokenization on data?