I would like to ask about how to finetune distillbart on gigaword and cnn dailymail with the starting checkpoint distilbart-cnn-12-6. I did use the gigaword dataset provided by tensorflow but it replaces numbers by this character: "#", as a result, my summaries have # instead of numbers, is it normal that it has those # ? Also is it really possible to finetune distillbart from the checkpoint distilbart-cnn-12-6 with cnn daily mail?
import os
os.environ['PYTHONPATH'] += ":/content/transformers/examples"
%cd "/content/transformers/examples"
!python /content/transformers/examples/seq2seq/finetune.py \
--learning_rate=3e-5 \
--fp16 \
--gpus 1 \
--do_train \
--do_predict \
--n_val 1000 \
--val_check_interval 0.1 \
--sortish_sampler \
--data_dir '/content/dataset' \
--train_batch_size=4 \
--eval_batch_size=4 \
--output_dir=distilbart_1300k_1400k \
--num_train_epochs 1 \
--model_name_or_path /content/transformers/examples/distilbart_1200k_1300k/best_tfmr
here is the link to gigaword: https://www.tensorflow.org/datasets/catalog/gigaword and here is the link to cnn dailymail: https://www.tensorflow.org/datasets/catalog/cnn_dailymail
For the code I followed the instruction of fine tuning distillabart here: https://github.com/Hildweig/transformers/tree/master/examples/seq2seq
For the outputs with gigawords I get something like this: "foreign exchange rates in hong kong sept. ## #### ; china 's defense minister says he 's ready to work with yugoslavia 's foreign minister on iraq 's role in iraq with bc-me-gen iraq"