I have a file containing BIO / IOB Tagged Text like this:
I got this from Wikineural and despite having the extension .conllu I think that it probably is not because these files have a different content. I need to store each column somewhere (idk what could be better) in order to have IDs, Words and Tags separated. With those data I have to do NER using HMM (Viterbi as decoding).
I'm new on python and NLP so don't go hard on me!! - I tried to store those data that way (Note that between the various sentences there are empty lines that I want to skip):
ids = []
words = []
tags = []
with open('en/train.conllu', encoding="utf8") as fileObj:
lines = [line.strip() for line in fileObj if line.strip()]
for line in lines:
row = line.split()
ids.append(row[0])
words.append(row[1])
tags.append(row[2])
but I feel like it's a bad way, any suggestion? Is there any particular data structure that I should use considering the final goal (NER)?
I have also another question NER related: since I have to use HMM - BIO/IOB tag already led me in a situation where I can start from learning (by counting TAG -> TAG and TAG -> WORD probability), right? I mean, those data are ready to be analyzed for NER purpose.
