Feature Columns Embedding lookup

Viewed 4278

I have been working with the datasets and feature_columns in tensorflow(https://developers.googleblog.com/2017/11/introducing-tensorflow-feature-columns.html). I see they have categorical features and a way to create embedding features from categorical features. But when working on nlp tasks, how do we create a single embedding lookup?

For eg: Consider text classification task. Every data point would have a lot of textual columns but they would not be separate categories. How do we create and use a single embedding lookup for all these columns?

Below is an example of how I am currently using the embedding features. I am building a categorical feature for each column and using that for creating embedding. The problem would be that the embeddings for same word could be different for different columns.

def create_embedding_features(key, vocab_list=None, embedding_size=20):
    cat_feature = \
        tf.feature_column.categorical_column_with_vocabulary_list(
            key=key,
            vocabulary_list = vocab_list
            )
    embedding_feature = tf.feature_column.embedding_column(
            categorical_column = cat_feature,
            dimension = embedding_size
        )
    return embedding_feature

le_features_embd = [create_embedding_features(f, vocab_list=vocab_list)
                     for f in feature_keys]
2 Answers

You are passing the weights directly to the "embedding_column" layer, but it is expecting callable object.

Best way is to create an initializer class and pass it to the embedding_column like this

import tensorflow as tf

class embedding_initialize(tf.keras.initializers.Initializer):
  def __call__(self, shape, dtype=None, **kwargs):
    return tf.compat.v1.Variable(initial_value=[[1, 2], [2,3],[3,4]], dtype = dtype)


one_hot_layer = tf.feature_column.sequence_categorical_column_with_identity('text', num_buckets=3)
text_embedding = tf.feature_column.embedding_column(one_hot_layer,
                                      dimension=2, 
                                      initializer = embedding_initialize())
columns = [text_embedding]

features = {'text': tf.sparse.from_dense([[1, 2], [2, 1]])}

sequence_input_layer = tf.keras.experimental.SequenceFeatures(columns)
sequence_input, sequence_length = sequence_input_layer(features, training=True)
Related