I am using spacy and its model en_core_web_lg, to perform summarisation in python. The code is running perfectly and there is no error at all. Except that, I am trying to find a way of making sure that the en_core_web_lg doesn't keep downloading in an environment if it already has it. I have googled a lot to find a perfect solution for this which I will list below but none has gelled with what I am trying to achieve. This code will be packaged and will be used by multiple people and I want to make sure that if they run the code everytime, the en_core_web_lg doesn't download if it already exists. Below is the spacy excerpt of my code and the solutions I tried:
#Importing necessary Libraries
from heapq import nlargest
from string import punctuation
import nltk
import spacy
from spacy.cli.download import download
from spacy.lang.en.stop_words import STOP_WORDS
nltk.download('punkt')
download(model="en_core_web_lg")
nlp_g = spacy.load('en_core_web_lg') #downloads everytime the code is run even if the model is present in the environment
def spacy_summarize(text):
"""
Returns the summary for an input string text
Parameters:
:param text: Input String
:type text: str
Returns:
:return: The summary for the input text
:rtype: String
"""
nlp = nlp_g
doc= nlp(text)
word_frequencies={}
for word in doc:
if word.text.lower() not in [list(STOP_WORDS), punctuation]:
if word.text not in word_frequencies:
word_frequencies[word.text] = 1
else:
word_frequencies[word.text] += 1
max_frequency=max(word_frequencies.values())
for word in word_frequencies:
word_frequencies.copy()[word]=word_frequencies[word]/max_frequency
sentence_tokens= [sent for sent in doc.sents]
sentence_scores = {}
spacy_frequencies(word_frequencies, sentence_tokens, sentence_scores)
select_length=max(1,int(len(sentence_tokens)*0.05))
summary=nlargest(select_length, sentence_scores,key=sentence_scores.get)
final_summary=[word.text for word in summary]
summary=''.join(final_summary)
return summary
def spacy_frequencies(word_frequencies, sentence_tokens, sentence_scores):
"""
Child function for spacy function for calculating sentence scores
Parameters:
:param: word frequeny, sentence token and score which
is provided through the parent function
"""
for sent in sentence_tokens:
for word in sent:
if word.text.lower() in word_frequencies:
if sent not in sentence_scores:
sentence_scores[sent]=word_frequencies[word.text.lower()]
else:
sentence_scores[sent]+=word_frequencies[word.text.lower()]
Things Tried:
import sys
import subprocess
import pkg_resources
required = {'en_core_web_lg'}
installed = {pkg.key for pkg in pkg_resources.working_set}
missing = required - installed
if missing:
python = sys.executable
subprocess.check_call([python, '-m', 'spacy', 'download', *missing], stdout=subprocess.DEVNULL)
try:
nlp_lg = spacy.load("en_core_web_lg")
except ModuleNotFoundError:
download(model="en_core_web_lg")
nlp_lg = spacy.load("en_core_web_lg")
Both solutions didn't give a satisfactory result and the package was downloaded again and I would appreciate if someone could help me with this? Thank you so much!