I have a very large dataset (~3mil) of addresses that I need geocoded, and I am using multiprocessing to speed it up. However, I've been using map, and it's eating up all my memory, causing other problems. I'd instead like to use imap, but it's not working as I'd expected. Here is my code:
import tqdm
import multiprocessing
import geocoder
def geocoding_mp(address):
pbar_gc.update(1)
return geocoder.osm(address, url = 'http://localhost:7070/').latlng
def mp(geocode_worker,df,num_proc):
pool = multiprocessing.Pool(processes = num_proc)
return pool.map(geocode_worker, df, chunksize = int(len(df)/num_proc))
tweets_usa_locations = ['Austin, TX','San Francisco, CA','Seattle, WA','Portland, ME'] # 3 million of these
if __name__ == '__main__':
pbar_gc = tqdm.tqdm(total = len(tweets_usa))
lat_long = mp(geocoding_mp,tweets_usa_locations,12)
When I do this, I get a nice list of coordinates that is the same length as my original list of addresses. When I change map to imap, and keep everything exactly the same, I get an iterator object, but when I try to open it using
for chunk in lat_long:
print(chunk)
It only prints 12 (the number of processes I used) coordinates, as opposed to 12 objects that each contain 3mil/12 coordinates. What's going on here? What else do I have to change when using imap to get an iterator that contains all of the locations?