There are many packages that have started to provide their own APIs to do the text preprocessing, however, each one has its own subtle differences.
tf.keras.preprocessing.text.Tokenizer() is implemented by Keras and is supported by Tensorflow as a high-level API.
tfds.features.text.Tokenizer() is developed and maintained by tensorflow itself.
Both have its own way of doing encoding the tokens. Which you can make out with the example below.
import tensorflow as tf
from tensorflow.keras.preprocessing.text import Tokenizer
import tensorflow_datasets as tfds
Let's take some sample data and see the encoded output for both the API's:
text_data = ["4. Kurt Betschart - Bruno Risi ( Switzerland ) 22",
"Israel approves Arafat 's flight to West Bank .",
"Moreau takes bronze medal as faster losing semifinalist .",
"W D L G / F G / A P",
"-- Helsinki newsroom +358 - 0 - 680 50 248",
"M'bishi Gas sets terms on 7-year straight ."]
First, let's the result for tf.keras.Tokenizer() :
tf_keras_tokenizer = Tokenizer()
tf_keras_tokenizer.fit_on_texts(text_data)
tf_keras_encoded = tf_keras_tokenizer.texts_to_sequences(text_data)
tf_keras_encoded = pad_sequences(tf_keras_encoded, padding="post")
For the first sentence in our input data, the result will be:
tf_keras_encoded[0]
array([2, 3, 4, 5, 6, 7, 8, 0], dtype=int32)
If we look at the word to index mapping.
tf_keras_tokenizer.index_word
{1: 'g',
2: '4',
3: 'kurt',
4: 'betschart',
5: 'bruno',
6: 'risi',
7: 'switzerland',
8: '22',
9: 'israel',
10: 'approves',
11: 'arafat',
12: "'s",
13: 'flight',
14: 'to',
15: 'west',
16: 'bank',
17: 'moreau',
18: 'takes',
19: 'bronze',
20: 'medal',
21: 'as',
22: 'faster',
23: 'losing',
24: 'semifinalist',
25: 'w',
26: 'd',
27: 'l',
28: 'f',
29: 'a',
30: 'p',
31: 'helsinki',
32: 'newsroom',
33: '358',
34: '0',
35: '680',
36: '50',
37: '248',
38: "m'bishi",
39: 'gas',
40: 'sets',
41: 'terms',
42: 'on',
43: '7',
44: 'year',
45: 'straight'}
Now let's try tfds.features.text.Tokenizer():
text_vocabulary_set = set()
for text in text_data:
text_tokens = tfds_tokenizer.tokenize(text)
text_vocabulary_set.update(text_tokens)
tfds_text_encoder = tfds.features.text.TokenTextEncoder(text_vocabulary_set, tokenizer=tfds_tokenizer)
For the first sentence in our input data, the result will be:
tfds_text_encoder.encode(text_data[0])
[35, 19, 44, 38, 32, 2, 14]
If we look at the word to index mapping(notice that index starts from 0 in this).
tfds_text_encoder._token_to_id
{'0': 0,
'22': 13,
'248': 17,
'358': 23,
'4': 34,
'50': 9,
'680': 6,
'7': 26,
'A': 19,
'Arafat': 39,
'Bank': 35,
'Betschart': 43,
'Bruno': 37,
'D': 15,
'F': 20,
'G': 28,
'Gas': 29,
'Helsinki': 38,
'Israel': 3,
'Kurt': 18,
'L': 44,
'M': 5,
'Moreau': 22,
'P': 10,
'Risi': 31,
'Switzerland': 1,
'W': 30,
'West': 33,
'approves': 4,
'as': 7,
'bishi': 2,
'bronze': 12,
'faster': 8,
'flight': 27,
'losing': 42,
'medal': 32,
'newsroom': 11,
'on': 25,
's': 24,
'semifinalist': 40,
'sets': 36,
'straight': 45,
'takes': 41,
'terms': 16,
'to': 14,
'year': 21}
You can see that the encoding difference in both the results along with that both the API's provides some hyperparameters which can be used and altered based on the requirement.