Reshaping Layer into Skew-Symmetric Matrix in Keras

Viewed 209

Say I have a keras model (for example)

layers_NE<-keras_model_sequential()
layers_NE %>% layer_dense(units=Height,
                           activation = "relu",
                           trainable=TRUE,
                           input_shape = 4,
                           bias_initializer = "random_normal") 
          %>% layer_dense(units = (d^2),
                           activation = "linear",
                           trainable = TRUE,
                           bias_initializer = "random_normal")

I want to reshape the last layer to a skew-symmetric matrix output for example like this c(a,b,c)-> c(c(a,b),c(b,c)) (here c(a,b,c) is notation for the output of my network)

So far I've tried this:

layers_NE %>%layer_reshape(input_shape = (d^2),
                           target_shape = c(d,d)
                           )

the output is of the correct shape but it isn't symmetric. How can I make that happen?

2 Answers

More frugal version which supposed to do exactly what you want without extra neurons. Also had to introduce two transpositions just to apply k_gather to correct axis, because of the way how k_gather is exposed to R (in python you could just pass axis=1 as argument to tf.gather):

Height <- 10
d <- 7
layers_NE<-keras_model_sequential()
layers_NE %>% layer_dense(units=Height,
                          activation = "relu",
                          trainable=TRUE,
                          input_shape = 4,
                          bias_initializer = "random_normal") 
%>% layer_dense(units = (d * (d+1) / 2),
                activation = "linear",
                trainable = TRUE,
                bias_initializer = "random_normal")
%>% layer_lambda(f=function(x) {
    selector <- array(0, dim=c(d^2))
    ind <- 0 # zero-based indicies needed here
    for (i in 1:d) {
        for (j in i:d) {
            selector[(i-1) * d + j] <- ind
            selector[(j-1) * d + i] <- ind
            ind <- ind + 1
        }
    }
    t_ind <- k_constant(selector, dtype='int32')
    k_permute_dimensions(k_gather(k_permute_dimensions(x, pattern=c(2,1)), t_ind), pattern=c(2,1))
})
%>% layer_reshape(input_shape = (d^2),
                  target_shape = c(d,d))

Upd. here is the similar piece, but in Python - the only important difference is that tf.gather can be called with axis=1:

def select_symmetric(x, d):
    selector = np.zeros(d*d, dtype=np.int32)
    ind = 0
    for i in range(d):
        for j in range(i, d):
            selector[i * d + j] = ind
            selector[j * d + i] = ind
            ind += 1
    t_ind = tf.constant(selector, dtype=tf.int32)
    return tf.gather(x, t_ind, axis=1)

Height = 10
d = 7
model = tf.keras.models.Sequential([
    tf.keras.layers.Dense(Height, 'relu', input_shape=(4,)),
    tf.keras.layers.Dense(d * (d + 1) // 2, 'linear'),
    tf.keras.layers.Lambda(select_symmetric, arguments={'d': d}),
    tf.keras.layers.Reshape(target_shape=(d, d)),
])

I see 2 ways to achieve that. The easier one would be to do exactly as you did initially - introduce at some point a layer with 2-dimensional square output per example (d,d) which is not symmetric and then make it symmetric by adding it to its own transposed version. It might look like the following:

layers_NE<-keras_model_sequential()
layers_NE %>% layer_dense(units=Height,
                           activation = "relu",
                           trainable=TRUE,
                           input_shape = 4,
                           bias_initializer = "random_normal") 
          %>% layer_dense(units = (d^2),
                           activation = "linear",
                           trainable = TRUE,
                           bias_initializer = "random_normal")
          %>%layer_reshape(input_shape = (d^2),
                           target_shape = c(d,d)
                           )
          %>% layer_lambda(f=function(x) {
                             (x + k_permute_dimensions(x, pattern=c(1,3,2))) * 0.5
                           })

After you add the model with its own transposed version the results will be symmetric (no real need to average here I guess). There is a bit of an redundancy in this solution since you have to actually train d^2 units instead of d(d+1)/2. Other than that it should be ok.

The second - a bit more frugal solution would be to actually create d(d+1)/2 units and put them into (d,d) shape in a way that non-diagonal elements are "duplicated". I believe what you are looking for is to create e.g. a lambda layer using k_gather function. But the only saving would be that you train less neurons on one of the layers.

Related