empty space after stopwords removal and lemmatisation

Viewed 57

The text looks like this before processing

0   [It's, good, for, beginners]                        positive
1   [I, recommend, this, starter, Ukulele, kit., I...   positive

After preprocessing with stopword removal and lemmatisation

nlp = spacy.load('en', disable=['ner', 'parser']) # disabling Named Entity Recognition for speed

def cleaning(doc):
    txt = [token.lemma_ for token in doc if not token.is_stop]
    if len(txt) > 2:
        return ' '.join(txt)
brief_cleaning = (re.sub("[^A-Za-z']+", ' ', str(row)).lower() for row in df3['reviewText'])

txt = [cleaning(doc) for doc in nlp.pipe(brief_cleaning, batch_size=5000, n_threads=-1)]

the result came like this

0   ' good ' ' ' ' beginner '                       positive
1   ' ' ' recommend ' ' ' ' starter ' ' ukulele ... positive

As you can see, there are lots of ' ' in the result, what caused this? I'm assuming it's the return ' '.join(txt) and re.sub("[^A-Za-z']+", ' ' that caused it, but if I removed the space or use return (txt), it simply won't remove any stopword or carry out lemmatisation.

Will these empty space cause troubles, or are they necessary, because I'm doing bigram and word2vec afterwards.

How can I fix it and have the result returned as ' recommend ' ' starter ' ' ukulele ' ' kit ' ' need ' ' learn ' ' ukulele '?

0 Answers
Related