Python RuntimeError: input sequence

Viewed 1201

I try to run NER in Indonesian Language

I've read some resources, they said that the BERT model has positional embeddings only for first 512 subtokens. So, the model can't work with longer sequences. I can't truncate text to 512 as there will be loss of information in that case.

There is also a model for long sequences, the name is 'Longformers'. But I can't use this model because the entity from deeppavlov is more complete than it.

Also I found sliding window approach at Pytorch error "RuntimeError: index out of range: Tried to access index 512 out of table with 511 rows", but I dont know how to implement it at NER usecase.

Could you please help how can I handle this?

The input is pandas.core.series.Series 'df'

2 Answers

Unlike other Transformers that have analytically computed position embeddings that can be potentially extended, position embeddings in BERT are trained jointly with the model, so there is no way to estimate what would BERT learn if it had longer outputs.

If you want to use BERT, you have no other choice than splitting your input text into segmented that are at most 512 sub-words long. From your example, it seems that your input texts consist of multiple sentences. You should definitely split the text into sentences before processing it with BERT (e.g., SpaCy can do a pretty decent sentence splitting), BERT is pre-trained on the sentence level. If the sentence splitting does not help (i.e., the sentences are still too long), you can come with some heuristic on how to further split the sentences (based on punctuation or something like that).

I seriously doubt that your data is of such nature that you can never ever segment it into smaller pieces. If they are, you should train a different model, something like LSTM-CRF. But I think that the harm from segmenting the text is much smaller than the benefit of using representation from BERT.

I am actually not sure if it makes a difference for Indonesian NER, but usually you don't want to simply split your texts. It is recommended to use a sliding window approach to preserve the context for the words on the edges of your 'simple' split.

I am also not a fan of frameworks which cover all the complexity from you because everything you want to do which is not implemented is complicated to achieve. Usually you would just modify one method and can leave to rest of your code to implement a sliding window split. It seems like that it is not that easy with deeppavlov because in this issue they recommended to check the length of the document before processing it with their pipeline (they call it chain).

Please have a look at the commented example below:

from bert_dp.tokenization import FullTokenizer
#we can use only 510 for our text because they are adding special tokens
maxtokens = 509
startOffset = 0
docStride = 200

#The location of your vocabulary file is probably different
#You can see the download location when you run:
#ner_model = build_model(configs.ner.ner_ontonotes_bert_mult, download=True)
tokenizer = FullTokenizer(vocab_file='/root/.deeppavlov/downloads/bert_models/multi_cased_L-12_H-768_A-12/vocab.txt', do_lower_case=False)

ok_doc = 'Limnologi mempelajari perairan di daratan Mamalogi mempelajari mamalia Mikrologi meneliti organisme mikroskopik dan interaksinya dengan kehidupan lainnya Mikologi mempelajari fungi Neurosains Neurobiologi mempelajari sistem saraf termasuk anatomi fisiologi dan patologinya Ornitologi mempelajari burung Paleontologi mempelajari fosil dan bukti Geografis kehidupan prasejarah Patologi Patobiologi atau patologi meneliti penyakit seperti penyebab proses ciri dan perkembangannya Parasitologi mempelajari parasit dan parasitisme Penelitian biomedis meneliti tubuh manusia yang sehat dan sakit Psikobiologi mempelajari dasar psikologi secara biologis Sosiobiologi mempelajari dasar sosiologi secara biologis Teknik biologis mempelajari biologi dari sudut pandang teknik dan lebih menekankan pada pengetahuan terapan Bidang ini terkait dengan bioteknologi Virologi mempelajari virus dan agen yang seperti virus Zoologi mempelajari hewan termasuk klasifikasi fisiologi perkembangan dan perilaku Bali adalah sebuah provinsi di Indonesia yang ibu kota provinsinya berNamaa Kota Denpasar Denpasar Bali juga merupakan salah satu pulau di Kepulauan Nusa Tenggara Di awal kemerdekaan Indonesia pulau ini termasuk dalam Provinsi Sunda Kecil yang beribu kota di Singaraja dan kini terbagi menjadi 3 provinsi Bali Nusa Tenggara Barat dan Nusa Tenggara Timur Selain terdiri dari Pulau Bali wilayah Provinsi Bali juga terdiri dari pulau-pulau yang lebih kecil di sekitarnya yaitu Nusa PenidaPulau Nusa Penida Nusa LembonganPulau Nusa Lembongan Nusa CeninganPulau Nusa Ceningan Pulau Serangan dan Pulau Menjangan Secara Geografis Bali terletak di antara Pulau Jawa dan Pulau Lombok Mayoritas puduk Bali adalah pemeluk agama Hindu Di dunia Bali terkenal sebagai tujuan pariwisata dengan keunikan berbagai hasil seni-budayanya khususnya bagi para wisatawan Jepang dan Australia Bali juga dikenal dengan julukan Pulau Dewata dan Pulau'
to_long_doc = ok_doc + ' Limnologi mempelajari ' + ok_doc 

docs = [ok_doc, to_long_doc]

def lenDocTokens(docTokens):
    length = sum([l for w,t,l in docTokens])
    return length

docsAfterSlidingWindow = []

for doc in docs:
    #tokenize your text
    docTokens = [(w, tokenizer.tokenize(w)) for w in doc.split()]
    docTokens = [(w, t, len(t)) for w,t in docTokens]
    print('doctTokens :{}'.format(lenDocTokens(docTokens)))

    while startOffset < lenDocTokens(docTokens):
        length = min(lenDocTokens(docTokens) - startOffset, maxtokens)
  
        slidingWindowDoc = []
        counter = 0

        #reconstructing the document
        for token in docTokens:
            counter += token[2]

            if (counter > startOffset+length):
                break
            elif (counter >= startOffset):
                b.append(token)
                slidingWindowDoc.append(token[0])

        docsAfterSlidingWindow.append(' '.join(slidingWindowDoc))
        #stop when the whole document is processed (document has less than 512
        #or the last document slice was processed)
        if startOffset + length == lenDocTokens(docTokens):
            break
        startOffset += min(length, docStride)
    startOffset = 0
print([len(tokenizer.tokenize(s)) for s in docsAfterSlidingWindow])

You can see that the second document was splitted into 3 sequences: Output:

doctTokens :447
doctTokens :902
[447, 508, 508, 503]
Related