So my goal is to train a BERT Model on wikipedia data that I derive right from Wikipedia. The contents that I scrape from the site look like this (example):
"(148975) 2001 XA255, provisional designation: 2001 XA255, is a dark minor planet in the outer Solar System, classified as centaur, approximately 38 kilometers (24 miles) in diameter. [...] \nFour temporary Neptune co-orbitals: (148975) 2001 XA255, (310071) 2010 KR59, (316179) 2010 EN65, and 2012 GX17 de la Fuente Marcos, C., & de la Fuente Marcos, R. 2012, Astronomy and Astrophysics, Volume 547, id.L2, 7 pp.\nIAU list of centaurs and scattered-disk objects\nIAU list of trans-neptunian objects\nAnother list of TNOs\n(148975) 2001 XA255 at the JPL Small-Body Database"
The Program that I wrote until now to clean up the strings looks like this:
import nltk
import string
import re
nltk.download('words')
words = set(nltk.corpus.words.words())
def clean_text(text):
'''
This function removes punctuation, words containing numbers
as well as making the whole text lower-case
'''
text = text.lower()
text = text.encode("ascii", errors="ignore").decode()
text = re.sub('[%s]' % re.escape(string.punctuation), '', text)
text = re.sub('\w*\d\w*', '', text)
text = re.sub('\n', '', text)
#text = " ".join(w for w in nltk.wordpunct_tokenize(text) if w.lower() in words or not w.isalpha())
text= text.split()
for word in text:
if len(word) <=1:
text.remove(word)
text = ' '.join([str(elem) for elem in text])
return text
By using this code the data is a bit cleaned up but still very messy. Output Any suggestions how to improve the function or should it be enough for the BERT model to run decently?