I'm looking for suggestions on how to approach a document classification problem. I will explain by means of example:
Problem statement
I have a collection of papers published by a university. I have another collection published by another university. And so on, for many universities.
When a new paper comes in, I'd like to determine which university probably published it.
My current approach
- For each university, build a dictionary of all terms from all papers with frequencies. Preprocess into terms and build a gensim
Dictionaryper university. - Build a "master" dictionary by merging all of the university-specific dictionaries. Iterate and perform
master_dictionary.merge_with(university_dictionary) - Treat each university dictionary as a document in a new corpus. Turn it into a BoW representation, and build a model from that. TfIdf/LSI/LDA.
- Perform a similarity match of the new paper (as BoW) against the model. Find the university dictionary document that matches closest.
My question
How is a problem like this tackled? (And is there a name for this?)
- I'm currently comparing a new document with a summarized document, one per university.
- I could compare a new document with every document across all universities. Then get the universities for the top matching documents using metadata.
- I've run across the Author-Topic Model but haven't looked into it, but that seems like this may be a good fit.
Any other ideas?