I was making a custom transformer model based on this tutorial for my dataset.
The overall shape of the dataset is like :
Encoder input : frame features of (Batch size, 150, 1024)
Decoder input : word sequence of (Batch size, 50), composed with [sos token, integer tokens]
Decoder output : word sequence of (Batch size, 50), composed with [integer tokens, eos token]
And each of them are named as "encoder_ipt_train", "decoder_ipt_train", "decoder_opt_train".
Here i used the code below for
1. Generating dataset batches
BATCH_SIZE = 32
BUFFER_SIZE = 16384
dataset = tf.data.Dataset.from_tensor_slices((
{
'inputs': encoder_ipt_train,
'dec_inputs': decoder_ipt_train
},
{
'outputs': decoder_opt_train
},
))
dataset = dataset.cache()
dataset = dataset.shuffle(BUFFER_SIZE)
dataset = dataset.batch(BATCH_SIZE)
dataset = dataset.prefetch(tf.data.experimental.AUTOTUNE)
2. In the training phase
EPOCHS = 20
train_step_signature = [
tf.TensorSpec(shape=(None, None), dtype=tf.int64),
tf.TensorSpec(shape=(None, None), dtype=tf.int64),
tf.TensorSpec(shape=(None, None), dtype=tf.int64),
]
@tf.function(input_signature=train_step_signature)
def train_step(enc_ipt, dec_ipt, dec_opt):
enc_padding_mask, combined_mask, dec_padding_mask = create_masks(enc_ipt, dec_ipt)
with tf.GradientTape() as tape:
predictions, _ = transformer(enc_ipt, dec_ipt, True,
enc_padding_mask, combined_mask, dec_padding_mask)
loss = loss_function(dec_opt, predictions)
gradients = tape.gradient(loss, transformer.trainable_variables)
optimizer.apply_gradients(zip(gradients, transformer.trainable_variables))
train_loss(loss)
train_accuracy(accuracy_function(dec_opt, predictions))
for epoch in range(EPOCHS):
start = time.time()
train_loss.reset_states()
train_accuracy.reset_states()
for (batch, (inp, tar)) in enumerate(dataset):
case 1 --> train_step(inp, tar)
case 2 --> train_step(inp['inputs'],inp['dec_inputs'], tar['outputs'])
if batch % 50 == 0:
print(f'Epoch {epoch + 1} Batch {batch} Loss {train_loss.result():.4f} Accuracy {train_accuracy.result():.4f}')
if (epoch + 1) % 5 == 0:
ckpt_save_path = ckpt_manager.save()
print(f'Saving checkpoint for epoch {epoch+1} at {ckpt_save_path}')
print(f'Epoch {epoch + 1} Loss {train_loss.result():.4f} Accuracy {train_accuracy.result():.4f}')
print(f'Time taken for 1 epoch: {time.time() - start:.2f} secs\n')
But this results a error like this, whether i use case 1 or case 2 in 2. :
The entire error message
TypeError: Expected any non-tensor type, got a tensor instead.
Now i'm doubting about my pipeline structure or the data type, but i have no idea how to change data type from 'dataset' in 1. Is there anyone who had similar problems like this or solved it?
and some additional questions :
First, I wonder if I'm using the tf.data.Dataset package in a right way. As far as i know, 'dataset' in 1 returns a dictionary of tensor batches, such that if i run the code below,
for (batch, (dict_1, dict_2)) in enumerate(dataset):
dict_1['inputs'] returns the batch of encoder_ipt_train with size of (32, 150, 1024). I wonder whether my understanding is correct.
Second, is there a better way(or a right way)to insert data batches to my training_step() function?
Sincerely, Seunghoon.