How do we use the output file generated after training a Stanford NER tagger using custom dataset?

Viewed 93

After following the steps in this Stanford NLP FAQ , I was able to generate a zip file of the model. But in the documentation they're using a TSV file to calculate the accuracy of prediction against already annotated file , but no documentation whatsoever is there as to how to test it against a new file!

Command used to generate the model was

 java -Xmx10240m -cp 'path_to_stanford-ner.jar' edu.stanford.nlp.ie.crf.CRFClassifier -prop austen.prop

where austen.prop is the properties which affect the training

Beginner in Java here , excuse if it's a silly question

1 Answers

The solution is to take the input file whichever you want to test against the model and to convert it into a TSV file which can be fed to the ner model by the following command

java -cp stanford-ner.jar edu.stanford.nlp.ie.crf.CRFClassifier -loadClassifier ner-model.ser.gz -testFile converted_to_tsv.tsv

Here's a small script to convert a file to TSV in python:

import json
import re
file = filepath
for line in open(file, mode="r",encoding = 'utf8'):
    regex = '[ ]'  
    with open('output.tsv','w+') as output_file:

        for line in list(filter(bool, file.splitlines())):

            for word in re.split(split_regex,line):
                print(word+"\tO")
                output_file.write(word+"\tO"+"\n")
Related