build siamese network via huggingface --- tokenize two sentences respectively using huggingface datasets and transformers along with tensorflow

Viewed 212

I'm currently building a siamese network with a pretrained Bert model which takes 'input_ids', 'token_type_ids' and 'attention_mask' as inputs from transformers. I've got a dataset structured as question1, question2, label, so I have to tokenize questions respectively.

def tokenize(ds):
    q1=datasets.Sequence(tokenizer(ds['question1'], padding='max_length', truncation=True, max_length=128))
    q2=datasets.Sequence(tokenizer(ds['question2'], padding='max_length', truncation=True, max_length=128))
    return {"q1":q1,"q2":q2}
dataset_tokenized = dataset.map(tokenize)

the process has gone halfway until it attempted to convert the result to pyarrow and raise an error:

ArrowInvalid: Could not convert Sequence(feature={'input_ids':[too_long_to_show], 'token_type_ids':[too_long_to_show],'attention_mask':[too_long_to_show]}, length=-1, id=None) with type Sequence: did not recognize Python value type when inferring an Arrow data type

In the official document in the section "Flatten" for datasets.Dataset, it seems that the dataset can take a Sequence as its feature.

I'd like to build the network like

class Siamese(Model, ABC):
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.transformer = TFBertModel.from_pretrained("hfl/chinese-bert-wwm")

    def call(self, inputs, training=None, mask=None):
        y1 = self.transformer(inputs[some_indices])
        y2 = self.transformer(inputs[some_indices])
        y1 = y1.get('last_hidden_state')
        y2 = y2.get('last_hidden_state')
        dist = tf.keras.losses.cosine_similarity(y1,y2)
        return dist

Question:

How can I arrange the dataset and the model to fit data?

1 Answers

Here is my solution. Though it's not elegant, it solved the problem. First, flatten the dataset:

dataset_tokenized = dataset_tokenized.flatten()

Then the columns went to something like (q1.input_ids, q1.token_type_ids, q1.attention_mask, q2.input_ids, q2.token_type_ids, q2.attention_mask) So we can do

features = {'q1':{'input_ids': tf_train_dataset['q1.input_ids'],
                  'token_type_eds':tf_train_dataset['q1.token_type_ids'],
                  'attention_mask':tf_train_dataset['q1.attention_mask']},
            'q2':{'input_ids':tf_train_dataset['q2.input_ids'],
                  'token_type_eds':tf_train_dataset['q2.token_type_ids'],
                  'attention_mask':tf_train_dataset['q2.attention_mask']}}

to make it to the desired structure.

Related