Which way is better to store data by column from file containing BIO/IOB tagged text?

Viewed 45

I have a file containing BIO / IOB Tagged Text like this:

BIO/IOB Tagged Text

I got this from Wikineural and despite having the extension .conllu I think that it probably is not because these files have a different content. I need to store each column somewhere (idk what could be better) in order to have IDs, Words and Tags separated. With those data I have to do NER using HMM (Viterbi as decoding).

I'm new on python and NLP so don't go hard on me!! - I tried to store those data that way (Note that between the various sentences there are empty lines that I want to skip):

ids = []
words = []
tags = []
with open('en/train.conllu', encoding="utf8") as fileObj:
    lines = [line.strip() for line in fileObj if line.strip()]
    for line in lines:
        row = line.split()
        ids.append(row[0])
        words.append(row[1])
        tags.append(row[2])

but I feel like it's a bad way, any suggestion? Is there any particular data structure that I should use considering the final goal (NER)?

I have also another question NER related: since I have to use HMM - BIO/IOB tag already led me in a situation where I can start from learning (by counting TAG -> TAG and TAG -> WORD probability), right? I mean, those data are ready to be analyzed for NER purpose.

0 Answers
Related