TensorFlow Tokenize by Unique CSV Row Rather than Unique Word in Row

Viewed 39

I'm trying to use tokenizer to classify free text into select categories.

For feature, I'm using this:

x_tokenizer = Tokenizer()
x_tokenizer.fit_on_texts(x)
x_train = x_tokenizer.texts_to_matrix(x_train, mode='count')
x_test = x_tokenizer.texts_to_matrix(x_test, mode='count')

and x_tokenizer.word_docs returns something like this:

defaultdict(<class 'int'>, {'name': 1, 'releasing': 1, 'one': 4, 'vehicle': 101, 'air': 3, 'vhel': 1, 'recently': 2})

This makes sense for feature, but I would like to use each row item for label. Right now, for label, I'm using the same code:

y_tokenizer = Tokenizer()
y_tokenizer.fit_on_texts(y)
y_train = y_tokenizer.texts_to_matrix(y_train, mode='count')
y_test = y_tokenizer.texts_to_matrix(y_test, mode='count')

and returns something like this:

defaultdict(<class 'int'>, {'a': 2, 'c': 2, 'language': 1, 'settings': 203, 'audio': 7, 'volume': 1})

but I would like to have this:

defaultdict(<class 'int'>, {'a/c': 2, 'language settings': 1, 'audio volume': 7})

so that every unique value in the label column would be represented as a unique token. How could I get this done?

Thank you in advance!

0 Answers
Related