Name of training and test data files in NLP (BioBERT GitHub repo)

Viewed 169

I'm reading the README.md file of the BioBERT GitHub repo:

Let $NER_DIR indicate a folder for a single NER dataset which contains train_dev.tsv, train.tsv, devel.tsv and test.tsv. Also, set $OUTPUT_DIR as a directory for NER outputs (trained models, test predictions, etc). For example, when fine-tuning on the NCBI disease corpus,

I don't understand what's the point of using 4 different files for training and testing. To me, only 2 are needed: train.tsv and test.tsv (and optionally valid.tsv).

Could someone explain the meaning of these files and why they are needed?

0 Answers
Related