I have a "simple" machine translation task where I have a sequence of vectors to be mapped to a word or two. (the vector is 258 dimensions)
For example:
[[1, ..., 2], [3, ..., 4]]=> "hello"[[1, ..., 2], [3, ..., 4], [5, ..., 6]]=> "hello world"
For the target field, I am using Field(eos_token="<eos>", is_target=True), which when batched does correctly give me a tensor with padding, in this case:
tensor([
[1, 1], # 1 is "hello"
[2, 3], # 2 is "world", 3 is <eos>
[3, 0], # 0 is <pad>
])
However, the src field is not being padded the same way, as it is sequential, but does not have a vocabulary (Field(use_vocab=False)).
When I read the src from the BucketIterator, in a batch of size > 1, I get:
Traceback (most recent call last):
File "train.py", line 50, in train
for b, batch in enumerate(train_iter):File "/torchtext/data/iterator.py", line 156, in iter
yield Batch(minibatch, self.dataset, self.device)File "/torchtext/data/batch.py", line 34, in init
setattr(self, name, field.process(batch, device=device))File "/torchtext/data/field.py", line 237, in process
tensor = self.numericalize(padded, device=device)File "/torchtext/data/field.py", line 359, in numericalize
var = torch.tensor(arr, dtype=self.dtype, device=device)ValueError: expected sequence of length 258 at dim 2 (got 5)
What I want to get is a tensor:
tensor([
[[1, ..., 2], [1, ..., 2]],
[[3, ..., 4], [3, ..., 4]],
[[5, ..., 6], [0, ..., 0]],
[[0, ..., 0], [0, ..., 0]],
])
What I think I might have but don't know how to confirm is:
tensor([
[[1, ..., 2], [1, ..., 2]],
[[3, ..., 4], [3, ..., 4]],
[[5, ..., 6], 0],
[0, 0],
])