NER ORG search for a company returns the word "company" instead of its name

Viewed 41

I'm working on an NLP/NER script using transformers/BERT and I'm having an issue extracting the name of a company from a set of texts.

In all the texts the script will be used on, the company's name will be presented like this:

"COMPANY NAME: the company's name is XXX" or "NAME: the company's name is XXX"

this is my code:

def get_company_info(text, tokenizer_1, model_1, tokenizer_2, model_2):
    company_info = {"name": None}
    try:
        start_company_index = re.search('name', text, re.I).span()[0]
        info = NLP_2(
            text[start_company_index:start_company_index+100], tokenizer_2, model_2)
        for data in info:
            if data['entity_group'] == 'ORG':
                company_info['name'] = data['word']
                break
    except:
        pass

However, the BERT script returns the word "company" since it finds it in the text and assumes correctly that it is the subject I'm looking for but I want to extract the name of the company instead.

Is there a simple way to avoid this or do I have to fine-tune the model?

I'm using regex to delimit the field of the search but I cannot simply use re.search("company") to start the search after the word company, because sometimes there will be 2 consecutive mentions of the word.

1 Answers

You cannot avoid that from a model-perspective unfortunately.

You have to:

  1. Either, as suggested in your comments, perform some regex parsing or try to use another type of logic to eliminate paragraph titles (if it suits you).
  2. Retrain on your own dataset, giving many examples of such situations like above, so that the network would eventually learn to distinguish and not detect COMPANY_NAME or other similar examples.

For (1) things can get complicated since what you gave here is just one instance where the network fails, it may very well be the case that as you see more documents, you discover more error-prone situations - the post-processing becomes more and more difficult.

For (2) you can actually start predicting on your new data, clean it, and start creating a cleaned dataset on your own for finetuning purposes.

One other approach would be to search for specific pre-trained BERT/similar models which are pre-trained on specific corpuses. For example, SciBERT is a pre-trained language model on scientific text, and, given that you would presumably work with scientific texts, it would have a better performance than a basic BERT. However, I do not know if you will find a model catered exactly for your needs as in the above example.

Related