PyTorch Dataset Field for Sequence of Vectors (No Vocabulary)

Viewed 376

I have a "simple" machine translation task where I have a sequence of vectors to be mapped to a word or two. (the vector is 258 dimensions)

For example:

  • [[1, ..., 2], [3, ..., 4]] => "hello"
  • [[1, ..., 2], [3, ..., 4], [5, ..., 6]] => "hello world"

For the target field, I am using Field(eos_token="<eos>", is_target=True), which when batched does correctly give me a tensor with padding, in this case:

tensor([
  [1, 1], # 1 is "hello"
  [2, 3], # 2 is "world", 3 is <eos>
  [3, 0], # 0 is <pad>
])

However, the src field is not being padded the same way, as it is sequential, but does not have a vocabulary (Field(use_vocab=False)).

When I read the src from the BucketIterator, in a batch of size > 1, I get:

Traceback (most recent call last):

File "train.py", line 50, in train

for b, batch in enumerate(train_iter):

File "/torchtext/data/iterator.py", line 156, in iter

yield Batch(minibatch, self.dataset, self.device)

File "/torchtext/data/batch.py", line 34, in init

setattr(self, name, field.process(batch, device=device))

File "/torchtext/data/field.py", line 237, in process

tensor = self.numericalize(padded, device=device)

File "/torchtext/data/field.py", line 359, in numericalize

var = torch.tensor(arr, dtype=self.dtype, device=device)

ValueError: expected sequence of length 258 at dim 2 (got 5)

What I want to get is a tensor:

tensor([
  [[1, ..., 2], [1, ..., 2]],
  [[3, ..., 4], [3, ..., 4]],
  [[5, ..., 6], [0, ..., 0]],
  [[0, ..., 0], [0, ..., 0]],
])

What I think I might have but don't know how to confirm is:

tensor([
  [[1, ..., 2], [1, ..., 2]],
  [[3, ..., 4], [3, ..., 4]],
  [[5, ..., 6], 0],
  [0, 0],
])
0 Answers
Related