Multiprocessing in NLTK

Viewed 45

I am trying to figure out a way to process a huge personal corpora in the following way and it's taking well beyond hours to process. I have a good enough CPU to speed it up manifold but I can't figure out the right way to apply multiprocessing to this code effectively. The code is as follows -:

#Import the COHA Corpus using NLTK

import nltk
from nltk.corpus import PlaintextCorpusReader
from nltk import word_tokenize
import os
from nltk import FreqDist
corpus_root = os.getcwd() + "/"
file_ids = ".*.txt"
#Tag
corpus = PlaintextCorpusReader(corpus_root, file_ids)
#tokenize and calculate overall freqdist
from multiprocessing import process
from collections import Counter
lst = []
for file in corpus.fileids():
    fdist = (corpus.words(file))
    lst.append(fdist)
    mlist = [item for sublist in lst for item in sublist]
    corpi = ' '.join([str(item) for item in mlist])
    corpora = nltk.word_tokenize(corpi)
    fdist = nltk.FreqDist(corpora)
    final = fdist.most_common(600)
    print(final)

Any help will be appreciated.

0 Answers
Related