Threadpool Executor still blocks in Python

Viewed 211

I am trying to make sure that I understand how to do non-blocking I/O in Python correctly. In looking at example from here I am still a little confused:

import concurrent.futures
import urllib.request

URLS = ['http://www.foxnews.com/',
   'http://www.cnn.com/',
   'http://europe.wsj.com/',
   'http://www.bbc.co.uk/',
   'http://some-made-up-domain.com/']

def load_url(url, timeout):
   with urllib.request.urlopen(url, timeout = timeout) as conn:
   return conn.read()

with concurrent.futures.ThreadPoolExecutor(max_workers = 5) as executor:
   future_to_url = {executor.submit(load_url, url, 60): url for url in URLS}
   for future in concurrent.futures.as_completed(future_to_url):
       url = future_to_url[future]
       try:
          data = future.result()
       except Exception as exc:
          print('%r generated an exception: %s' % (url, exc))
       else:
          print('%r page is %d bytes' % (url, len(data)))

This block of code downloads the URLs in parallel, but let's say I have something that needs updates without blocking (like printing a status update every .2 seconds.

Where does that code go?

It seems like even though the URL downloads are not blocking each other, the overall task of downloading all the URLs is blocking the rest of the program while they complete downloading.

0 Answers
Related