I am trying to train a model for entity recognition on resumes. More specifically, I am trying to train a model to recognize education, professional experience, skills, etc.. on resumes. I am using a dataset of resumes I found online that is already formatted in a way that a spacy 'ner' model would recognize. But the dataset is in English, and I need French data. At some point, I will probably build the dataset manually, but for now I am going to settle for translating the dataset I already have. For example, let's manufacture a datapoint:
[['I went to New York ', {entity : [11,19, Location], [3, 7, verb]}]]. The numbers represent the position of the first and last character. So 'New York' is a location.
So the issue here is that the translation will shift, change, the position of the entities that are important for us. So then my question is : Is there a better way to do this ?