Issues with spacy model en_core_web_lg : how to prevent the package from downloading every time the code is run

Viewed 92

I am using spacy and its model en_core_web_lg, to perform summarisation in python. The code is running perfectly and there is no error at all. Except that, I am trying to find a way of making sure that the en_core_web_lg doesn't keep downloading in an environment if it already has it. I have googled a lot to find a perfect solution for this which I will list below but none has gelled with what I am trying to achieve. This code will be packaged and will be used by multiple people and I want to make sure that if they run the code everytime, the en_core_web_lg doesn't download if it already exists. Below is the spacy excerpt of my code and the solutions I tried:

#Importing necessary Libraries
from heapq import nlargest
from string import punctuation
import nltk
import spacy
from spacy.cli.download import download
from spacy.lang.en.stop_words import STOP_WORDS

nltk.download('punkt')
download(model="en_core_web_lg")
nlp_g = spacy.load('en_core_web_lg') #downloads everytime the code is run even if the model is present in the environment

def spacy_summarize(text):
    """
    Returns the summary for an input string text

            Parameters:
                :param text: Input String
                :type text: str

            Returns:
                :return: The summary for the input text
                :rtype: String

    """
    nlp = nlp_g
    doc= nlp(text)
    word_frequencies={}
    for word in doc:
        if word.text.lower() not in [list(STOP_WORDS), punctuation]:
            if word.text not in word_frequencies:
                word_frequencies[word.text] = 1
            else:
                word_frequencies[word.text] += 1
    max_frequency=max(word_frequencies.values())
    for word in word_frequencies:
        word_frequencies.copy()[word]=word_frequencies[word]/max_frequency
    sentence_tokens= [sent for sent in doc.sents]
    sentence_scores = {}
    spacy_frequencies(word_frequencies, sentence_tokens, sentence_scores)
    select_length=max(1,int(len(sentence_tokens)*0.05))
    summary=nlargest(select_length, sentence_scores,key=sentence_scores.get)
    final_summary=[word.text for word in summary]
    summary=''.join(final_summary)
    return summary

def spacy_frequencies(word_frequencies, sentence_tokens, sentence_scores):
    """
    Child function for spacy function for calculating sentence scores
            Parameters:
                    :param: word frequeny, sentence token and score which
                        is provided through the parent function

    """
    for sent in sentence_tokens:
        for word in sent:
            if word.text.lower() in word_frequencies:
                if sent not in sentence_scores:
                    sentence_scores[sent]=word_frequencies[word.text.lower()]
                else:
                    sentence_scores[sent]+=word_frequencies[word.text.lower()]



Things Tried:

import sys
import subprocess
import pkg_resources

required = {'en_core_web_lg'}
installed = {pkg.key for pkg in pkg_resources.working_set}
missing = required - installed

if missing:
  python = sys.executable
  subprocess.check_call([python, '-m', 'spacy', 'download', *missing], stdout=subprocess.DEVNULL)

try:
  nlp_lg = spacy.load("en_core_web_lg")
except ModuleNotFoundError:
  download(model="en_core_web_lg")
  nlp_lg = spacy.load("en_core_web_lg")


Both solutions didn't give a satisfactory result and the package was downloaded again and I would appreciate if someone could help me with this? Thank you so much!

1 Answers

spaCy doesn't automatically download models at all, so this must be a bug with your code that checks if the model is already installed.

Looking at this code:

try:
  nlp_lg = spacy.load("en_core_web_lg")
except ModuleNotFoundError:
  download(model="en_core_web_lg")
  nlp_lg = spacy.load("en_core_web_lg")

The issue is that if the model is not installed this is an OSError, not a ModuleNoteFoundError. First you need to fix that.

This approach seems like it should work, except loading models in the same process you installed them in doesn't work very reliably - the list of installed packages is not updated while Python is running. So even after fixing the above issue, it may not work as intended.

I would recommend either:

  1. Download the model to a known directory, extract it there, and load it from a path instead of just the model name
  2. Check the output of pip list to see if the model is installed, and install it if not
Related