Handling missing categorical levels in the out of sample data using num_oov_buckets option of categorical_column_with_vocabulary_list function

Viewed 170

I have a categorical feature called feature 1 that has 3 possible values (value1, value2, value3). I want to create an embedding feature column of feature 1 for the model with taking into consideration encountering a new value in the out of sample dataset.

So the first step was to use the tf.feature_column.categorical_column_with_vocabulary_list function to create a feature columns type that can be accepted by the tf.feature_column.embedding_column function.

Inside the tf.feature_column.categorical_column_with_vocabulary_list function we can specify in the attributes how to handle the out-of-vocabulary values presented in the out of sample dataset. Therefore, I used the num_oov_buckets option and set it to 1.

In second place, I used the tf.feature_column.embedding_column function to embed the 4 categorical values(value1, value2, value, oov bucket added in the previous step) into two-dimensional space.

I am wondering how the model attributed those weights to the new bucket despite seeing only 3 values (value1, value2, value3) in the train set.

enter image description here

0 Answers
Related