String cleaning/preprocessing for BERT

Viewed 178

So my goal is to train a BERT Model on wikipedia data that I derive right from Wikipedia. The contents that I scrape from the site look like this (example):


"(148975) 2001 XA255, provisional designation: 2001 XA255, is a dark minor planet in the outer Solar System, classified as centaur, approximately 38 kilometers (24 miles) in diameter. [...] \nFour temporary Neptune co-orbitals: (148975) 2001 XA255, (310071) 2010 KR59, (316179) 2010 EN65, and 2012 GX17 de la Fuente Marcos, C., & de la Fuente Marcos, R. 2012, Astronomy and Astrophysics, Volume 547, id.L2, 7 pp.\nIAU list of centaurs and scattered-disk objects\nIAU list of trans-neptunian objects\nAnother list of TNOs\n(148975) 2001 XA255 at the JPL Small-Body Database"


The Program that I wrote until now to clean up the strings looks like this:

import nltk
import string
import re
nltk.download('words')
words = set(nltk.corpus.words.words())

def clean_text(text):
    '''
    This function removes punctuation, words containing numbers
    as well as making the whole text lower-case
    '''
    text = text.lower()
    text = text.encode("ascii", errors="ignore").decode()
    text = re.sub('[%s]' % re.escape(string.punctuation), '', text)
    text = re.sub('\w*\d\w*', '', text)
    text = re.sub('\n', '', text)
    #text = " ".join(w for w in nltk.wordpunct_tokenize(text) if w.lower() in words or not w.isalpha())
    text= text.split()
    for word in text:
        if len(word) <=1:
            text.remove(word)
    text = ' '.join([str(elem) for elem in text])    
    return text

By using this code the data is a bit cleaned up but still very messy. Output Any suggestions how to improve the function or should it be enough for the BERT model to run decently?

0 Answers
Related