Scraping large number of pages (full website) with Python and Selenium

Viewed 64

I'm trying to scrap the full web page with python and selenium. I'm providing the website URL (i.e. https://example.com) which then will scrap the homepage and then it will continue doing this recursively for all subpages by getting the links from the first page. I'm facing serious performance issues as the number of pages can exceed 500 and the selenium will stop in the middle of the execution.

FYI I'm using Selenium as I need to scan JS-based web pages (where content is being generated by JavaScript). Are there any best practices on how to Scrap the full web page content more efficiently? or other similar libraries?

1 Answers

Here is a solution with concurrent.futures.ThreadPoolExecutor module,

The code below will run multiple selenium scraping processes parallelly


def product_parser(product_links):

    # your selenium code

    driver = webdriver.Chrome(
        executable_path=CHROME_DRIVER_PATH,
        options=chrome_options,
    )
    ........
    ........


# create a list of 10 lists with product links
list_chunked_product_links = [[], [], [],......[]]  

# an example, it will create a list of 10 lists consists of max 1000 elements each
# [product_links[x:x + 1000] for x in range(0, len(product_links) - 2, 1000)]


# run 10 drivers concurrently
def run_all(product_links):
    with concurrent.futures.ThreadPoolExecutor(max_workers=10) as executor:
        executor.map(product_parser, product_links)


def main_func():
    start_time = time.time()
    t_start = time.localtime()
    run_all(list_chunked_product_links)
    duration = time.time() - start_time
    print(f"start time : {time.strftime('%H:%M:%S', t_start)}, \n End Time : {time.strftime('%H:%M:%S')} ")
    print(f"Scraped {len(product_links)} in {duration} seconds")


Related