I have a tool that generates a large number of files (ranging from hundreds of thousands to few millions) whenever it runs. All these files can be read independently of each other. I need to parse them and summarize the information.
Dummy example of generated files:
File1:
NAME=John AGE=25 ADDRESS=123 Fake St
NAME=Jane AGE=25 ADDRESS=234 Fake St
File2:
NAME=Dan AGE=30 ADDRESS=123 Fake St
NAME=Lisa AGE=30 ADDRESS=234 Fake St
Summary - counts how many times an address appeared across all files:
123 Fake St - 2
234 Fake St - 2
I want to use parallelization to read them, so multiprocessing or asyncio come to mind (I/O intensive operations). I plan to do the following operations in a single unit/function that will be called in parallel for each file:
- Open the file, go line by line
- Populate a unique dict containing the information provided by this file specifically
- Close the file
Once I am done reading all the files in parallel, and have one dict per file, I can now loop over each dict and summarize as needed.
The reason I think I need this two step process, is I can't have multiple parallel calls to that function directly summarize and write to a common summary dict. That will mess things up.
But that means I will consume a large amount of memory (due to holding those many hundreds of thousands to millions of dicts in memory).
What would be a good way to get the best of both worlds - runtime and memory consumption - to meet this objective?