NLP: Find the most similar Corpus (not Document)

Viewed 45

I'm looking for suggestions on how to approach a document classification problem. I will explain by means of example:

Problem statement

I have a collection of papers published by a university. I have another collection published by another university. And so on, for many universities.

When a new paper comes in, I'd like to determine which university probably published it.

My current approach

  1. For each university, build a dictionary of all terms from all papers with frequencies. Preprocess into terms and build a gensim Dictionary per university.
  2. Build a "master" dictionary by merging all of the university-specific dictionaries. Iterate and perform master_dictionary.merge_with(university_dictionary)
  3. Treat each university dictionary as a document in a new corpus. Turn it into a BoW representation, and build a model from that. TfIdf/LSI/LDA.
  4. Perform a similarity match of the new paper (as BoW) against the model. Find the university dictionary document that matches closest.

My question

How is a problem like this tackled? (And is there a name for this?)

  • I'm currently comparing a new document with a summarized document, one per university.
  • I could compare a new document with every document across all universities. Then get the universities for the top matching documents using metadata.
  • I've run across the Author-Topic Model but haven't looked into it, but that seems like this may be a good fit.

Any other ideas?

0 Answers
Related