I am currently working on my CustomDownloadMiddleware. This is mainly due to a certain need for control over downloading the webpage. My CustomDownloadMiddleware looks something as below
class MySeleniumDownloadMiddleware:
"""Scrapy middleware handling the requests using selenium"""
def __init__(self, driver):
self.driver = driver
self.cookies = self.driver.get_cookies()
@classmethod
def from_crawler(cls, crawler):
"""Initialize the middleware with the crawler settings"""
driver = init_chromium(crawler.settings.get('SELENIUM_HOSTNAME'))
login(driver, crawler.settings.get('MY_CREDENTIAL'))
middleware = cls(driver=driver)
crawler.signals.connect(middleware.spider_closed, signals.spider_closed)
return middleware
def process_request(self, request, spider):
"""Process a request using the selenium driver if applicable"""
try:
self.driver.get(request.url)
except WebDriverException:
self.driver = init_chromium(spider.settings.get('SELENIUM_HOSTNAME'))
recover_cookie_to_driver(self.driver, self.cookies)
self.driver.get(request.url)
body = str.encode(self.driver.page_source)
# Expose the driver via the "meta" attribute
request.meta.update({'driver': self.driver})
return HtmlResponse(
self.driver.current_url,
body=body,
encoding='utf-8',
request=request
)
def spider_closed(self):
"""Shutdown the driver when spider is closed"""
try:
self.driver.quit()
except WebDriverException:
pass
After opening an issue in github, I had found that the issue is that custom download middleware are required to handle the download delay by itself, and the recommendation are by using slot that can be found in this file. However, I had fail to find any relevant example of the slot being used in the download middleware.
Also, it is worth to mention that my code are heavily inspired by scrapy-selenium, and had found an issue which provide a potential solution for download delay, however when I use time.sleep(second) solution, the pipeline and spider's parse function are not being processed concurrently, which I deem the solution not viable.
Is there a proper way of implementing slot into download middleware so it respect download delay, as well as it will not hinder the ability of both pipeline and scraping being processed concurrently? If yes, what is it?