Comparison of tf.keras.preprocessing.text.Tokenizer() and tfds.features.text.Tokenizer()

Viewed 1657

As some background, I've been looking more and more into NLP and text-processing lately. I am much more familiar with Computer Vision. I understand the idea of Tokenization completely.

My confusion stems from the various implementations of the Tokenizer class that can be found within the Tensorflow ecosystem.

There is a Tokenizer class found within Tensorflow Datasets (tfds) as well as one found within Tensorflow proper: tfds.features.text.Tokenizer() & tf.keras.preprocessing.text.Tokenizer() respectively.

I looked into the source code (linked below) but was unable to glean any useful insights


The tl;dr question here is: Which library do you use for what? And what are the benefits of one library over the other?


NOTE

I was following along with the Tensorflow In Practice Specialization as well as this tutorial. The TF in Practice Specialization uses the tf.Keras.preprocessing.text.Tokenizer() implementation and the text loading tutorial uses tfds.features.text.Tokenizer()

1 Answers

There are many packages that have started to provide their own APIs to do the text preprocessing, however, each one has its own subtle differences.

tf.keras.preprocessing.text.Tokenizer() is implemented by Keras and is supported by Tensorflow as a high-level API.

tfds.features.text.Tokenizer() is developed and maintained by tensorflow itself.

Both have its own way of doing encoding the tokens. Which you can make out with the example below.

import tensorflow as tf
from tensorflow.keras.preprocessing.text import Tokenizer
import tensorflow_datasets as tfds  

Let's take some sample data and see the encoded output for both the API's:

text_data = ["4. Kurt Betschart - Bruno Risi ( Switzerland ) 22",
            "Israel approves Arafat 's flight to West Bank .",
            "Moreau takes bronze medal as faster losing semifinalist .",
            "W D L G / F G / A P",
            "-- Helsinki newsroom +358 - 0 - 680 50 248",
            "M'bishi Gas sets terms on 7-year straight ."]  

First, let's the result for tf.keras.Tokenizer() :

tf_keras_tokenizer = Tokenizer()
tf_keras_tokenizer.fit_on_texts(text_data)
tf_keras_encoded = tf_keras_tokenizer.texts_to_sequences(text_data)
tf_keras_encoded = pad_sequences(tf_keras_encoded, padding="post") 

For the first sentence in our input data, the result will be:

tf_keras_encoded[0]  

array([2, 3, 4, 5, 6, 7, 8, 0], dtype=int32)

If we look at the word to index mapping.

tf_keras_tokenizer.index_word  


{1: 'g',
 2: '4',
 3: 'kurt',
 4: 'betschart',
 5: 'bruno',
 6: 'risi',
 7: 'switzerland',
 8: '22',
 9: 'israel',
 10: 'approves',
 11: 'arafat',
 12: "'s",
 13: 'flight',
 14: 'to',
 15: 'west',
 16: 'bank',
 17: 'moreau',
 18: 'takes',
 19: 'bronze',
 20: 'medal',
 21: 'as',
 22: 'faster',
 23: 'losing',
 24: 'semifinalist',
 25: 'w',
 26: 'd',
 27: 'l',
 28: 'f',
 29: 'a',
 30: 'p',
 31: 'helsinki',
 32: 'newsroom',
 33: '358',
 34: '0',
 35: '680',
 36: '50',
 37: '248',
 38: "m'bishi",
 39: 'gas',
 40: 'sets',
 41: 'terms',
 42: 'on',
 43: '7',
 44: 'year',
 45: 'straight'}  

Now let's try tfds.features.text.Tokenizer():

text_vocabulary_set = set()
for text in text_data:
    text_tokens = tfds_tokenizer.tokenize(text)
    text_vocabulary_set.update(text_tokens) 

tfds_text_encoder = tfds.features.text.TokenTextEncoder(text_vocabulary_set, tokenizer=tfds_tokenizer)  

For the first sentence in our input data, the result will be:

tfds_text_encoder.encode(text_data[0]) 

[35, 19, 44, 38, 32, 2, 14]

If we look at the word to index mapping(notice that index starts from 0 in this).

tfds_text_encoder._token_to_id  

{'0': 0,
 '22': 13,
 '248': 17,
 '358': 23,
 '4': 34,
 '50': 9,
 '680': 6,
 '7': 26,
 'A': 19,
 'Arafat': 39,
 'Bank': 35,
 'Betschart': 43,
 'Bruno': 37,
 'D': 15,
 'F': 20,
 'G': 28,
 'Gas': 29,
 'Helsinki': 38,
 'Israel': 3,
 'Kurt': 18,
 'L': 44,
 'M': 5,
 'Moreau': 22,
 'P': 10,
 'Risi': 31,
 'Switzerland': 1,
 'W': 30,
 'West': 33,
 'approves': 4,
 'as': 7,
 'bishi': 2,
 'bronze': 12,
 'faster': 8,
 'flight': 27,
 'losing': 42,
 'medal': 32,
 'newsroom': 11,
 'on': 25,
 's': 24,
 'semifinalist': 40,
 'sets': 36,
 'straight': 45,
 'takes': 41,
 'terms': 16,
 'to': 14,
 'year': 21}  

You can see that the encoding difference in both the results along with that both the API's provides some hyperparameters which can be used and altered based on the requirement.

Related