Scrapy starts 8 crawlers all at the same time. 3 of them are on the same IP and so we set CONCURRENT_REQUESTS = 1 this is per spider, but then we also need the 3 spiders that access the same server (ip) to wait for each other.
How and where can we create some kind of queue mechanism in the BaseSpider() that if one of the 3 spiders was already started then the other 2 need to wait.
Possible solution A:
- Suppose the 3 spiders are called
Heuy, Dewey and Louie - create
wait_for list variablethat is set on spider level. ForHeuyit is empty, forDeweyit is set to[Heuy, Louie]and forLouieit is set to[Heuy, Dewey] - In this example all 3 would start and run into some if/else logic in BaseSpider init where we a) check for the
wait_for list variablethen access it and see if any of the other spiders are running, if so then we pause() the current spider. ForHeuythis would mean I wat for nobody and start.Dewey and Louiehave await_forand becauseHeuyis running the both will pause. - In the close spider logic of BaseSpider we do the same check again, only this time we access the all waiting spiders list, get their
wait_forlist, check if any of thewait_forspiders is still running (is it? when close spider is still being accessed ...) and unpause() the spider (which isLouieorDewey). - Then when the 2nd spider finishes it does the same check again and starts the 3rd spider
Possible solution B:
- Suppose the 3 spiders are called
Heuy, Dewey and Louie - create
crawl_pool_queueon BaseSpider level. - create
crawl_pool=ducksvariable that is set on spider level forHeuy, Dewey and Louie. - In this example all 3 would start and run into some if/else logic in BaseSpider init where we check for the existance of
crawl_pool - If it exists then we access all running spiders and see if already 1 in the same
crawl_poolare running. If not then start ... if yes then push tocrawl_pool_queue - The paused spiders are stored in crawl_pool_queue using a dict containing a list per key for waiting spiders ---- this allows us to have more than 1 queue.
- If a spider finishes(closes) then we try to access
crawl_poolif it does not exist then there is no spider waiting; if it does exist then we accesscrawl_pool_queueand find the next spider name that needs to run, pop it from the list and start it - Etc Etc
To me solution B sound like the best solution. Although I am not sure.
So my question is:
- what is the best way to create a scalable crawler_queue where spiders wait for each other when we cannot access the same domain at the same time?
- what is the best place to store this logic? Would it be init for pausing the spiders?
- what is the best place to store this logic? Would it be closed() for unpausing the spiders? -- I mean is the spider then really closed already?
- can we create some code that can be used by the community for this to share
Sample code for solution B
class BaseShirtsSpider(Spider):
# dict to store pools and lists of waiting crawlers
crawl_pool_queue = {}
# default there is no crawl pool, not ture if I have to define it as False
crawl_pool = False
# Maybe a method even before init_login (ideally before even making any request, or at the minimum the 1st loginpage)
def init_login(self, meta=None):
# Only check logic if crawl_pool is defined
if self.crawl_pool:
# How can I retrieve all spiders and their state running/paused
for spider in crawler.spiders():
# If we find a RUNNING spider and it has the same crawl pool then pause it() + add to crawl_pool_queue
if spider.crawler.engine.state = 1 and spider.crawler.engine.crawl_pool = self.crawl_pool:
crawl_pool_queue[crawl_pool].append(spider.crawler.name)
spider.crawler.engine.pause()
break
def closed(self, reason):
# Only check logic if crawl_pool is defined
if self.crawl_pool:
# How can I retrieve all spiders and their state running/paused
for spider in crawler.spiders():
# If we find a WAITING spider and it has the same crawl pool then unpause it() + remote from crawl_pool_queue
if spider.crawler.engine.state = 0 and spider.crawler.engine.crawl_pool = self.crawl_pool:
crawl_pool_queue[crawl_pool].remove(spider.crawler.name)
spider.crawler.engine.unpause()
break
super().closed(reason)
class HueySpider(BaseShirtsSpider):
name = 'huey'
crawl_pool = 'ducks'
start_urls = ['samedomain.com']
class DeweySpider(BaseShirtsSpider):
name = 'dewey'
crawl_pool = 'ducks'
start_urls = ['samedomain.com']
class LouieSpider(BaseShirtsSpider):
name = 'louie'
crawl_pool = 'ducks'
start_urls = ['samedomain.com']
class DonaldSpider(BaseShirtsSpider):
name = 'donald'
start_urls = ['someotherdomain.com']