SUMY Text Summarizer fails to summarize and returns original text

Viewed 753
LANGUAGE = "english"
stemmer = Stemmer(LANGUAGE)

def get_luhn_summary(text):
        summ = list()
    
        parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))
        summarizer = LuhnSummarizer()
        summarizer.stop_words = get_stop_words(LANGUAGE)
    
        for sentence in summarizer(parser.document,10):
            summ.append(str(sentence))
        return summ

summaryA_luhn = get_luhn_summary(textA)

Always returns the original string. I am confused cause I am following the documentation to the t

1 Answers

The summarization is done by sentence count.

import nltk
from sumy.parsers.plaintext import PlaintextParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.luhn import LuhnSummarizer as Summarizer
from sumy.nlp.stemmers import Stemmer
from sumy.utils import get_stop_words

LANGUAGE = "english"
SENTENCES_COUNT = 2
nltk.download('punkt')

parser = PlaintextParser.from_file("document.txt", Tokenizer(LANGUAGE))

stemmer = Stemmer(LANGUAGE)

summarizer = Summarizer(stemmer)
summarizer.stop_words = get_stop_words(LANGUAGE)

for sentence in summarizer(parser.document, SENTENCES_COUNT):
    print(sentence)

The following will read sentences from file name document.txt and based on SENTENCES_COUNT it will summarize based on the number of sentences you specify.

So if document.txt has 10 sentences and you set SENTENCES_COUNT = 2 you will get a summarization of two sentences.

You can also simply swap out:

parser = PlaintextParser.from_file("document.txt", Tokenizer(LANGUAGE))

with:

text = "This is the string to parse. Hopefully it will be more than one sentence. Like so!"    
parser = PlaintextParser.from_string(text, Tokenizer(LANGUAGE))

If you what to parse from string instead of a file.

Related