How to finetune distillbart for abstractive summarization using Gigaword or Cnn dailymail?

Viewed 481

I would like to ask about how to finetune distillbart on gigaword and cnn dailymail with the starting checkpoint distilbart-cnn-12-6. I did use the gigaword dataset provided by tensorflow but it replaces numbers by this character: "#", as a result, my summaries have # instead of numbers, is it normal that it has those # ? Also is it really possible to finetune distillbart from the checkpoint distilbart-cnn-12-6 with cnn daily mail?

import os
os.environ['PYTHONPATH'] += ":/content/transformers/examples"
%cd "/content/transformers/examples"

!python /content/transformers/examples/seq2seq/finetune.py \
    --learning_rate=3e-5 \
    --fp16 \
    --gpus 1 \
    --do_train \
    --do_predict \
    --n_val 1000 \
    --val_check_interval 0.1 \
    --sortish_sampler \
    --data_dir '/content/dataset' \
    --train_batch_size=4 \
    --eval_batch_size=4 \
    --output_dir=distilbart_1300k_1400k \
    --num_train_epochs 1 \
    --model_name_or_path /content/transformers/examples/distilbart_1200k_1300k/best_tfmr

here is the link to gigaword: https://www.tensorflow.org/datasets/catalog/gigaword and here is the link to cnn dailymail: https://www.tensorflow.org/datasets/catalog/cnn_dailymail

For the code I followed the instruction of fine tuning distillabart here: https://github.com/Hildweig/transformers/tree/master/examples/seq2seq

For the outputs with gigawords I get something like this: "foreign exchange rates in hong kong sept. ## #### ; china 's defense minister says he 's ready to work with yugoslavia 's foreign minister on iraq 's role in iraq with bc-me-gen iraq"

0 Answers
Related