More efficient HashMap (Dictionary) for Python for use in big data

Viewed 1353

I'm creating a program that counts the occurrences of strings in a huge file. For this I have used the python dictionary with the strings as keys and the counts as the values.

The program works fine for smaller files of up to 10000 strings. But when I test it out on my actual file ~ 2-3 mil strings, my program starts to slow down when it reaches the 50% mark to about 50% of it's original speed.

I suspect this is because the built in dictionary was not meant to handle such large volumes of data and is getting a lot more collisions. I would like to know if there is any efficient way of solving this problem. I was looking for alternative hashmap implementations or even making a list of hashmaps (it slowed it down further).

Details:

  • The strings are not known beforehand.
  • The strings' lengths range is about 10 - 200.
  • There are many strings that only occur once (and will be discarded at the end)
  • I have already implemented concurrency to speed it up.
  • It takes about 1 hour to complete one file
    • I do other calculations too, while that takes up time, it does not slow down on smaller filesizes. So I suspect it's a hashmap or memory issue.
  • I have plenty of memory, when running it only takes up 8GB of 32GB.
1 Answers
Related