ThreadpoolExecotor missing some rows in python

Viewed 86

I was experimenting with the multithreading and could not understand why it is skipping many rows? please help me improve my code can i use shutodown fn to improve it?

import concurrent.futures
import requests
import json
import time

URLList = []
for i in range(100):
    URLList.append("https://randomuser.me/api/")
def getName(urlApi):
        print(requests.get(urlApi).json()["results"][0]["name"])
def main():
    with concurrent.futures.ThreadPoolExecutor(max_workers=10) as executor:
        executor.submit(getName)
        executor.map(getName , URLList)

start_time = time.time()
main()
midtime = time.time() - start_time

print(f"--- {midtime} seconds ---")

Output is some random rows not 100 rows

{'title': 'Miss', 'first': 'Lily', 'last': 'Johnson'}
{'title': 'Mr', 'first': 'میلاد', 'last': 'سلطانی نژاد'}
{'title': 'Ms', 'first': 'Celina', 'last': 'Henry'}
{'title': 'Mr', 'first': 'آرمین', 'last': 'سالاری'}
{'title': 'Ms', 'first': 'Stans', 'last': 'Ordelman'}
{'title': 'Mr', 'first': 'Matthäus', 'last': 'Hilger'}
{'title': 'Ms', 'first': 'Marta', 'last': 'Flores'}
{'title': 'Mr', 'first': 'Alex', 'last': 'Gutierrez'}
{'title': 'Mr', 'first': 'Hugo', 'last': 'Walker'}
{'title': 'Mr', 'first': 'Vincent', 'last': 'Campbell'}
{'title': 'Mr', 'first': 'آریا', 'last': 'صدر'}
{'title': 'Mr', 'first': 'Lode', 'last': 'Mulderij'}
{'title': 'Miss', 'first': 'Manuela', 'last': 'Rubio'}
{'title': 'Mr', 'first': 'Jacob', 'last': 'Wang'}
{'title': 'Mr', 'first': 'رضا', 'last': 'کامروا'}
{'title': 'Mr', 'first': 'سهیل', 'last': 'مرادی'}
{'title': 'Miss', 'first': 'Kristin', 'last': 'Roberts'}
{'title': 'Mr', 'first': 'Allen', 'last': 'Weaver'}
{'title': 'Miss', 'first': 'Cathy', 'last': 'Johnston'}
{'title': 'Mr', 'first': 'Isaías', 'last': 'da Paz'}
{'title': 'Mr', 'first': 'Felix', 'last': 'Lam'}
{'title': 'Mr', 'first': 'Julio', 'last': 'Chambers'}
{'title': 'Mr', 'first': 'Felix', 'last': 'Johansen'}
{'title': 'Mr', 'first': 'Ruben', 'last': 'Mora'}
{'title': 'Mr', 'first': 'Marco', 'last': 'Diaz'}
{'title': 'Ms', 'first': 'Yvonne', 'last': 'Sims'}
{'title': 'Mr', 'first': 'Conrad', 'last': 'Owren'}
{'title': 'Mr', 'first': 'آرسین', 'last': 'گلشن'}
{'title': 'Mr', 'first': 'Aitor', 'last': 'Santos'}
{'title': 'Ms', 'first': 'Sally', 'last': 'Bell'}
{'title': 'Mr', 'first': 'Valentin', 'last': 'Marquez'}
{'title': 'Mr', 'first': 'Jacob', 'last': 'Mortensen'}
{'title': 'Monsieur', 'first': 'Nicolas', 'last': 'Lucas'}

Please Assist me.

1 Answers

I am able to get all 100 results when using the following snippet:

def main():
    with concurrent.futures.ThreadPoolExecutor(max_workers=10) as executor:
        results = [executor.submit(getName , x) for x in URLList]
    print(len(results)) 

However, when I try to actually look at the results:

[result.result() for result in results]

I get:

JSONDecodeError: Expecting value: line 1 column 1 (char 0)

This is because some of the requests do not actually come back successful.

So let's try changing getName a little bit so we see what happens when we fail:

def getName(urlApi):
  request = requests.get(urlApi)
  try:
    return request.json()["results"][0]["name"]
  except:
    return request.text

And for we get for some rows is:

<html>
<head>
<title>Uh oh, something bad happened</title>
</head>
<body align='center'>
<h1 align='center'>Uh oh, something bad happened</h1>
<a class="twitter-timeline" data-width="450" data-height="700" data-theme="light" href="https://twitter.com/randomapi">Tweets by randomapi</a> <script async src="//platform.twitter.com/widgets.js" charset="utf-8"></script>
</body>
</html>

If you try to reduce the number of workers, you will see that more requests come back successful, which suggests that you are facing some rate limiting. In any case, as far as multithreading is concerned, everything is working correctly and you are indeed sending a hundred requests and receiving a hundred responses. It's just that some of them aren't exactly the JSON responses you were hoping to get.

When setting max_workers to 1 I was able to get all 100 responses come back successfully. This way the requests are sent sequentially but asynchronously which is still parallel and therefore faster than simply sending them over one by one, waiting for each response before sending the next request.

With max_workers up to 4, I was still able to get 100% of the requests successfully but at max_workers=5 about half of them failed. You have to try and see what rate works best - you can always resend the failed requests but it's better to find the highest possible rate with no failures at all.

Related