Sparse feature vector sequence as input to LSTM

Viewed 249

I need to build an LSTM model on a my input data which is sparse vector sequence. Each sample is of the format: [v_1, v_2,...,v_t] where each v_t is the sparse feature vector at time t with format [i_1, i_2, ..., i_n] where i_j is the index of the feature with 1 as value (everything else is 0). Normally the number of non-zero features are about 0.001 of the total features, so data is pretty sparse. What I do now is that for each batch I convert the sparse data into dense numpy matrix and I pass to the LSTM directly (I can't convert the whole data into dense format because it won't fit in the memory). I believe I could benefit from using embedding. I was thinking of getting embedding for each index and then sum them up but the embedding layer requires that each time step has same number of features which is not the case here. I think pytorch would be more flexible to handle this but I prefer to use keras/tensorflow as much as possible. Thanks.

1 Answers

Turned out that RaggedTensors are what I was looking for. They are very powerful in representing sparse data. Here's an example that worked for me. I still need to figure out how they work with DataGenerator. Based on my experience it takes more time to load data into the tensor but once Input tensor is loaded the epochs are faster than using the dense tensors.

data = tf.ragged.constant([ [[940, 203, 668, 387, 790, 320, 939, 185],[315, 515, 791, 181, 939, 787]], 
                             [[564, 205], [820, 180, 993, 739]] ])

input_dim = 1000
X = Input(shape=[None, None], dtype=tf.int64, ragged=True)
l1 = Embedding(input_dim, 16)(X)
l2 = tf.reduce_sum(l1, axis=2) #To calculate the dense feature vector for each timestep.
l3 = LSTM(32, use_bias=False) (l2)
l4 = Dense(32)(l3)
l5 = Activation(tf.nn.relu)(l4)
output = tf.keras.layers.Dense(1)(l5)    

model = Model(X, output)
print(model.summary())
model.compile(loss='binary_crossentropy', optimizer='rmsprop')
print(model(data))   
Related