I would like to verify whether the code I have created is an appropriate and effective way to reuse a aiohttp.ClientSession() object.
from asyncio import tasks
import aiohttp
import asyncio
import time
websites = ['http://corndog.io/', 'https://onesquareminesweeper.com/', 'https://checkboxolympics.com/', 'https://binarypiano.com/', 'https://alwaysjudgeabookbyitscover.com/', 'https://cant-not-tweet-this.com/', 'https://cursoreffects.com/', 'http://eelslap.com/', 'https://smashthewalls.com/', 'https://thatsthefinger.com/']
# function to regroup a larger list of tasks to "smaller groups"
def feeder(a_list, unit):
list_length = len(a_list)
groups = []
a = 0
z = unit
while a < list_length:
groups.append(a_list[a:z])
a += unit
z += unit
return groups
# function to create a list of co-routine object from each "smaller group" for the "gather()" function
def create_tasks(session, job_slice):
co_routine_objects = []
for i in job_slice:
co_routine_objects.append(session.get(i, ssl = False))
return co_routine_objects
async def main(job_list):
async with aiohttp.ClientSession() as session:
for job_slice in job_list:
tasks = create_tasks(session, job_slice)
responses = await asyncio.gather(*tasks)
for item in responses:
data.append(await item.text())
print(len(data))
job_list = feeder(websites, 3)
data = []
loop = asyncio.get_event_loop()
loop.run_until_complete(main(job_list))
What I tried to achieve:
1 - Take a list of website URLS
2 - Regroup the URLS within a list of lists
3 - Create a synchronous function which has the role to return a list of co-routine objects built from an item within the list of lists created in the previous step
4 - Create an asynchronous function which has the role to CREATE AND KEEP OPEN A REUSABLE AIOHTTP CLIENT SESSION OBJECT then post a list of GET requests returned by the "create tasks" function all at once
5 - Take the next "group" of links from the regrouped list and feed the GET requests through the SAME CLIENT SESSION THAT WAS CREATED IN STEP 4
I am in the process of learning to work with the "aiohttp" and the "async" module and while the code above is functioning I am not sure if it is functioning as I intended (see above) or if there is a better, or more efficient/faster way to implement/structure the code.
Could someone please confirm that:
- only one session object is used
- the session object created is kept open and reused throughout the execution of the code and not "opened and closed" with each iteration
Also if you see a way to restructure the sample code that would allow to obtain data in a faster more efficient way from the URLS using AIOHTTP/ASYNCIO, please highlight it.
Many thanks.