Visualizing network of sentences in Textrank

Viewed 91

I'm using the Textrank method explained here to get the summary of the text. Is there a way to plot the output of the textrank_sentences like a network of all the textrank_ids connected to each other?

library(textrank)
data(joboffer)

library(udpipe)
tagger <- udpipe_load_model(tagger$file_model)
joboffer <- udpipe_annotate(tagger, job_rawtxt)
joboffer <- as.data.frame(joboffer)
joboffer$textrank_id <- unique_identifier(joboffer, c("doc_id","paragraph_id", "sentence_id"))
sentences <- unique(joboffer[, c("textrank_id", "sentence")])
terminology <- subset(joboffer, upos %in% c("NOUN", "ADJ"))
terminology <- terminology[, c("textrank_id", "lemma")]
tr <- textrank_sentences(data = sentences, terminology = terminology)
1 Answers

This question is rather old, but is a good question and deserves an answer.

Yes! textrank returns all the information that you need. Just look at the output of str(tr). Part of it says:

 $ sentences_dist:Classes ‘data.table’ and 'data.frame':        666 obs. of  3 variables:
  ..$ textrank_id_1: int [1:666] 1 1 1 1 1 1 1 1 1 1 ...
  ..$ textrank_id_2: int [1:666] 2 3 4 5 6 7 8 9 10 11 ...
  ..$ weight       : num [1:666] 0.1429 0.4167 0 0.0625 0 ...

This gives which sentences are connected in the form of a lower triangular matrix. Two sentences are connected if the weight of their connection is greater than zero. To visualize the graph, use the non-zero weights as an edgelist and build the graph.

Links = which(tr$sentences_dist$weight > 0)
EdgeList = cbind(tr$sentences_dist$textrank_id_1[Links],
        tr$sentences_dist$textrank_id_2[Links])
library(igraph)
SGraph1 = graph_from_edgelist(EdgeList, directed=FALSE)

set.seed(42)
plot(SGraph1)

Graph 1

We see that 11 of the nodes (sentences) are not connected to any other node. For example, sentences 15 and 36

tr$sentences$sentence[c(36,15)]
[1] "Contact:"                                                 
[2] "Integration of the models into the existing architecture."

But other other nodes do connect up, for example node 1 is connected to node 2.

tr$sentences$sentence[c(1,2)]
[1] "Statistical expert / data scientist / analytical developer"                           
[2] "BNOSAC (Belgium Network of Open Source Analytical Consultants), 
is a Belgium consultancy company specialized in data analysis and 
statistical consultancy using open source tools."

because those sentences share the (important) words "statistical", "data", and "analytical".

The singleton nodes take up a lot of space in the graph making the other nodes rather crowded. So I will also show the graph with those removed.

which(degree(SGraph1) == 0)
 [1]  4  7 15 20 21 23 25 26 29 30 36

SGraph2 = delete.vertices(SGraph1, which(degree(SGraph1) == 0))
set.seed(42)
plot(SGraph2)

Graph 2

That shows the relations between sentences somewhat better, but I expect that you can find a nicer layout for the graph that better shows the relations. However, that is not the thrust of the question and I leave it to you to make the graph pretty.

Related