I'm trying to use tokenizer to classify free text into select categories.
For feature, I'm using this:
x_tokenizer = Tokenizer()
x_tokenizer.fit_on_texts(x)
x_train = x_tokenizer.texts_to_matrix(x_train, mode='count')
x_test = x_tokenizer.texts_to_matrix(x_test, mode='count')
and x_tokenizer.word_docs returns something like this:
defaultdict(<class 'int'>, {'name': 1, 'releasing': 1, 'one': 4, 'vehicle': 101, 'air': 3, 'vhel': 1, 'recently': 2})
This makes sense for feature, but I would like to use each row item for label. Right now, for label, I'm using the same code:
y_tokenizer = Tokenizer()
y_tokenizer.fit_on_texts(y)
y_train = y_tokenizer.texts_to_matrix(y_train, mode='count')
y_test = y_tokenizer.texts_to_matrix(y_test, mode='count')
and returns something like this:
defaultdict(<class 'int'>, {'a': 2, 'c': 2, 'language': 1, 'settings': 203, 'audio': 7, 'volume': 1})
but I would like to have this:
defaultdict(<class 'int'>, {'a/c': 2, 'language settings': 1, 'audio volume': 7})
so that every unique value in the label column would be represented as a unique token. How could I get this done?
Thank you in advance!