Multiple requests with aiohttp, but seperate timeout per request

Viewed 2303

I have a huge list of URL's (around 40 million).

I wrote a script that scrapes this URL's with multithreading. But I need an extra solution, which should be economical in OS resources, so I've decided to develop the ASYNC version as well.

I've studied asyncio and aiohttp in Python for a week.

Below is the working code:

from pathlib import Path
import time
import asyncio
import aiohttp
import pypeln as pl
import async_timeout


# for calculating the total elapsed time
start = time.time()

successful_counter = 0

# files and folders
urlFile = open('url500.txt', 'r')



# list for holding processed url's so far
urlList = []


#######################
# crawler function start
#######################
async def crawling(line, session1):  # function wrapper for parallelizing the process
    # getting URL's from the file
    
    global successful_counter
    
    line = line.strip()  
    
    # try to establish a connection
    try:
        async with async_timeout.timeout(25):
            async with session1.get('http://' + line) as r1:
                x = r1.headers
                if ('audio' in x['Content-Type'] or 'video' in x['Content-Type']):
                    print("Url: " + line + " is a streaming website \n")
                    return  # stream website, skip this website

                # means we have established a connection and got the expected result
                if r1.status // 100 == 2:
                    #print("Returned 2** for the URL:", line)
                    
                    try:
                        text1 = await r1.text()
                        successful_counter += 1

                        '''
                        f1 = open('200/' + line + '.html', 'w')
                        f1.write(text1)
                        f1.close()
                        '''

                    except Exception as exc:
                        print(line + ": " + str(exc))
                        return
                    
                    urlList.append(line)
                    return
                                
                else:
                    return

    # some error occured
    except Exception as exc:
        print("Url: " + line + " created the error: \n" + str(exc))
        return                
            
        
#######################
# crawler function end
#######################


async def main(tempList):

    '''
    limit = 1000
    await pl.task.each(
            crawling, tempList, workers=limit,
        )
    '''
    conn = aiohttp.TCPConnector(limit=0)
    custom_header1 = {'User-agent': 'Mozilla/5.0 (X11; Linux i586; rv:31.0) Gecko/20100101 Firefox/74.0'}


    #'''
    async with aiohttp.ClientSession(headers=custom_header1, connector=conn) as session1:
        await asyncio.gather(*[asyncio.ensure_future(crawling(url, session1)) for url in tempList])
    #'''

    return


asyncio.run(main(urlFile))

print("total successful: ", successful_counter)

# for calculating the total elapsed time
end = time.time()
print("Total elapsed time in seconds:", end-start)

Here is the problem: when I don't put any timeouts, it works with no problems but takes too much time. I want to spend at most 25 seconds per request, if the website does not give me any response, I should skip that website, and move on.

So far, every method I've tried failed me. When I put a timeout of 25 seconds in somewhere, it always restricts the whole program, instead of the single request. So whether I have a file that has 500 URLs, or 1000000 URLs, it's always ending in 25 seconds.

I've tried wrapping the crawler function with async_timeout, using the built-in timeout of aiohttp library

async with session1.get('http://' + line, timeout=25)

Tried to create session inside the crawler function and put a timeout on the session (again using aiohttp's built-in methods).

Nothing worked... Probably I'm missing something huge, but I'm stuck for days, and ran out of options to try :D

2 Answers

As a starting point; I would recommend creating a small script to test the bare minimum to allow a get request to timeout without affecting other requests.

In the code below, the timeout is set to half a second. All the URLs are the same (stackoverflow.com) except for one, which points to localhost (which is used to test the timeout). Also if the url is stackoverflow.com, the code sleeps for 2 seconds (to show timeout).

import asyncio
import aiohttp
import json

test_url = "https://stackoverflow.com/"

def Logger(json_message):
    print(json.dumps(json_message))

async def get_data(url):
    Logger({"start": "get_data()", "url": url})
    if url is test_url: #This is a test to make "test url" sleep longer than the timeout.   
        await asyncio.sleep(2) 

    timeout = aiohttp.ClientTimeout(total=0.5) # TODO - timeout after half a second.
    try:
        async with aiohttp.ClientSession(timeout=timeout) as session:
            async with session.get(url) as results:            
                Logger({"finish": "get_data()", "url": url})
                return f"{ results.status } - {url}"
    except Exception as exc:
        Logger({"error": "get_data()", "url": url, "message": str(exc) })
        return f"fail - {url}"

async def main():
    urls = [test_url]*5 # create array of 5 urls
    urls[2] = "https://localhost:44344/" # Set third url to something that will timeout (after 0.5 sec).
    statements = [get_data(x) for x in urls]    
    Logger({"start": "gather()"})

    results = await asyncio.gather(*statements) 
    Logger({"finish": "gather()"})
    Logger({"results": ", ".join(results)})

if __name__ == '__main__':
    #asyncio.set_event_loop_policy(asyncio.WindowsSelectorEventLoopPolicy()) # Use this to stop "Event loop is closed" error on Windows - https://github.com/encode/httpx/issues/914
    asyncio.run(main())

Output:

{"start": "gather()"}
{"start": "get_data()", "url": "https://stackoverflow.com/"}
{"start": "get_data()", "url": "https://stackoverflow.com/"}
{"start": "get_data()", "url": "https://localhost:44344/"}
{"start": "get_data()", "url": "https://stackoverflow.com/"}
{"start": "get_data()", "url": "https://stackoverflow.com/"}
{"error": "get_data()", "url": "https://localhost:44344/", "message": ""}
{"finish": "get_data()", "url": "https://stackoverflow.com/"}
{"finish": "get_data()", "url": "https://stackoverflow.com/"}
{"finish": "get_data()", "url": "https://stackoverflow.com/"}
{"finish": "get_data()", "url": "https://stackoverflow.com/"}
{"finish": "gather()"}
{"results": "200 - https://stackoverflow.com/, 200 - https://stackoverflow.com/, fail - https://localhost:44344/, 200 - https://stackoverflow.com/, 200 - https://stackoverflow.com/"}

From the quickstart documentation there are 4 different parts of the timeout you can set: total, connect, sock_connect, and sock_read.

total

The maximal number of seconds for the whole operation including connection establishment, request sending and response reading.

connect

The maximal number of seconds for connection establishment of a new connection or for waiting for a free connection from a pool if pool connection limits are exceeded.

sock_connect

The maximal number of seconds for connecting to a peer for a new connection, not given from a pool.

sock_read

The maximal number of seconds allowed for period between reading a new data portion from a peer.

Setting the total timeout will probably mean some requests are having to wait for pool connections to free up, and hitting the timeout.

You could try

timeout = aiohttp.ClientTimeout(sock_connect=25)
async with aiohttp.ClientSession(timeout=timeout) as session:
    # Use session to perform requests...
Related