I am implementing a type of search (TF-IDF) in which every word has a score calculated that is proportional to all the documents being searched. I have 100GB of documents to search.
If I was working with 1GB of documents, I would use:
Dictionary<string, List<Document>>
..where string is the word and List<Document> is all documents, ranked in order, containing that word. This does not scale up. I am using a Dictionary<> because lookup time is O(1) (in theory).
My intended solution is a SQLServer database in which words are listed in a table, with the relevant List object stored serialized. My concern is that reading the DB and rebuilding to List<> each time will be very inefficient.
Am I going in the wrong direction here? What is a normal solution to working with huge dictionaries?