Page similarity calculation with Gensim

Viewed 66

there are some questions about this which I've studied, but I'm still not sure about a couple of things.
I want to split a page stream (concatenated PDFs) in to singular documents. So the trick is to find where one document ends and where the next one starts. So a PDF can have 1000 pages, and can consist out of 20 documents, each with different lengths.

That being said, one feature I want to introduce is 'page similarity' where page (p) has a similarity score for the page before it (p-1). So studying this problem leads me to a lot of examples using LDA and LSI models, but is this the way to go?

I have made a corpus with all the tokens, bigrams, trigrams from all 1000 pages. What is the best way to compare two pages with each other? I have looked had this example where an LSI model is used to compare a query with a whole corpus, but I can't figure out how to compare it with just one previous page/document. Any other ideas will be greatly appreciated!

texts = data_lemmatized      # --> all tokenized + filtered + bigrams + trigrams using gensim
dictionary = corpora.Dictionary(data_lemmatized)
corpus = [dictionary.doc2bow(text) for text in data_lemmatized]

lsi = models.LsiModel(corpus, id2word=dictionary, num_topics=2)
vec_bow = dictionary.doc2bow(data_lemmatized[1]).   #--> this is page 2, which I want to compare with data_lemmatized[0]
vec_lsi = lsi[vec_bow]

index = similarities.MatrixSimilarity(lsi[corpus])
sims = index[vec_lsi]  # -->this performs a similarity query against the corpus, but I want only 1 page
1 Answers

There are many different possible text-similarity calculations; which one is the "way to go" will depend on your data, project goals, & resources. Both LDA & LSI are reasonable things to try.

The Gensim models work on whatever 'documents' you give them - so if you've preprocessed your 1000-page PDF into 20 of your true documents, or 1000 separate pages, before training a topic model, those are also the units-of-text it will analyze.

Are you trying to compare the pages to infer document-boundaries? (Are you sure there's no other better hint of boundaries in the PDF? Are you sure all the desired document boundaries align with page-boundaries?)

Doing page-level comparisons might work for that, or might not it'd depend on how vividly the word/phrase usage changes from the end of one document to another.

You can generally think of many of the Gensim models as providing a summary vector for a text. The queries against all documents are useful for listing ranked matches, but if you want to do a simple pairwise calculation, you'd get the vectors for each text from the model individually, then use a direct calculation on the vectors (such as cosine-similarity). For example:

page1_bow = dictionary.doc2bow(data_lemmatized[0])
page1_lsi = lsi[page1_bow]
page2_bow = dictionary.doc2bow(data_lemmatized[1])
page2_lsi = lsi[page2_bow]
cossim = gensim.matutils.cossim(page1_lsi, page2_lsi)

(The gensim.mattutils.cossim() function is a convenience helper for calculating cosine-similarity between the sparse arrays many Gensim bag-of-words-fed models, including LSI. With other dense vectors, you might use the raw calculation for cosine-similarity, or the cosine-distance function offered by scipy.spatial.distance.cosine, or other methods.)

Related