I'm creating a program that counts the occurrences of strings in a huge file. For this I have used the python dictionary with the strings as keys and the counts as the values.
The program works fine for smaller files of up to 10000 strings. But when I test it out on my actual file ~ 2-3 mil strings, my program starts to slow down when it reaches the 50% mark to about 50% of it's original speed.
I suspect this is because the built in dictionary was not meant to handle such large volumes of data and is getting a lot more collisions. I would like to know if there is any efficient way of solving this problem. I was looking for alternative hashmap implementations or even making a list of hashmaps (it slowed it down further).
Details:
- The strings are not known beforehand.
- The strings' lengths range is about 10 - 200.
- There are many strings that only occur once (and will be discarded at the end)
- I have already implemented concurrency to speed it up.
- It takes about 1 hour to complete one file
- I do other calculations too, while that takes up time, it does not slow down on smaller filesizes. So I suspect it's a hashmap or memory issue.
- I have plenty of memory, when running it only takes up 8GB of 32GB.