Scrapy, create queues for spiders using a extendable queue pool mechanism

Viewed 74

Scrapy starts 8 crawlers all at the same time. 3 of them are on the same IP and so we set CONCURRENT_REQUESTS = 1 this is per spider, but then we also need the 3 spiders that access the same server (ip) to wait for each other.

How and where can we create some kind of queue mechanism in the BaseSpider() that if one of the 3 spiders was already started then the other 2 need to wait.

Possible solution A:

  • Suppose the 3 spiders are called Heuy, Dewey and Louie
  • create wait_for list variable that is set on spider level. For Heuy it is empty, for Dewey it is set to [Heuy, Louie] and for Louie it is set to [Heuy, Dewey]
  • In this example all 3 would start and run into some if/else logic in BaseSpider init where we a) check for the wait_for list variable then access it and see if any of the other spiders are running, if so then we pause() the current spider. For Heuy this would mean I wat for nobody and start. Dewey and Louie have a wait_for and because Heuy is running the both will pause.
  • In the close spider logic of BaseSpider we do the same check again, only this time we access the all waiting spiders list, get their wait_for list, check if any of the wait_for spiders is still running (is it? when close spider is still being accessed ...) and unpause() the spider (which is Louie or Dewey).
  • Then when the 2nd spider finishes it does the same check again and starts the 3rd spider

Possible solution B:

  • Suppose the 3 spiders are called Heuy, Dewey and Louie
  • create crawl_pool_queue on BaseSpider level.
  • create crawl_pool = ducks variable that is set on spider level for Heuy, Dewey and Louie.
  • In this example all 3 would start and run into some if/else logic in BaseSpider init where we check for the existance of crawl_pool
  • If it exists then we access all running spiders and see if already 1 in the same crawl_pool are running. If not then start ... if yes then push to crawl_pool_queue
  • The paused spiders are stored in crawl_pool_queue using a dict containing a list per key for waiting spiders ---- this allows us to have more than 1 queue.
  • If a spider finishes(closes) then we try to access crawl_pool if it does not exist then there is no spider waiting; if it does exist then we access crawl_pool_queue and find the next spider name that needs to run, pop it from the list and start it
  • Etc Etc

To me solution B sound like the best solution. Although I am not sure.

So my question is:

  • what is the best way to create a scalable crawler_queue where spiders wait for each other when we cannot access the same domain at the same time?
  • what is the best place to store this logic? Would it be init for pausing the spiders?
  • what is the best place to store this logic? Would it be closed() for unpausing the spiders? -- I mean is the spider then really closed already?
  • can we create some code that can be used by the community for this to share

Sample code for solution B

class BaseShirtsSpider(Spider):

    # dict to store pools and lists of waiting crawlers
    crawl_pool_queue = {}
    # default there is no crawl pool, not ture if I have to define it as False
    crawl_pool = False

    # Maybe a method even before init_login (ideally before even making any request, or at the minimum the 1st loginpage)
    def init_login(self, meta=None):
        # Only check logic if crawl_pool is defined
        if self.crawl_pool:
            # How can I retrieve all spiders and their state running/paused
            for spider in crawler.spiders():
                # If we find a RUNNING spider and it has the same crawl pool then pause it() + add to crawl_pool_queue
                if spider.crawler.engine.state = 1 and spider.crawler.engine.crawl_pool = self.crawl_pool:
                    crawl_pool_queue[crawl_pool].append(spider.crawler.name)
                    spider.crawler.engine.pause()
                    break

    def closed(self, reason):
        # Only check logic if crawl_pool is defined
        if self.crawl_pool:
            # How can I retrieve all spiders and their state running/paused
            for spider in crawler.spiders():
                # If we find a WAITING spider and it has the same crawl pool then unpause it() + remote from crawl_pool_queue
                if spider.crawler.engine.state = 0 and spider.crawler.engine.crawl_pool = self.crawl_pool:
                    crawl_pool_queue[crawl_pool].remove(spider.crawler.name)
                    spider.crawler.engine.unpause()
                    break
        super().closed(reason)

class HueySpider(BaseShirtsSpider):
    name = 'huey'
    crawl_pool = 'ducks'
    start_urls = ['samedomain.com']

class DeweySpider(BaseShirtsSpider):
    name = 'dewey'
    crawl_pool = 'ducks'
    start_urls = ['samedomain.com']

class LouieSpider(BaseShirtsSpider):
    name = 'louie'
    crawl_pool = 'ducks'
    start_urls = ['samedomain.com']

class DonaldSpider(BaseShirtsSpider):
    name = 'donald'
    start_urls = ['someotherdomain.com']
0 Answers
Related