Use scrapyd job id in scrapy pipelines

Viewed 952

I've implemented a web application that is triggering scrapy spiders using scrapyd API (web app and scrapyd are running on the same server).

My web application is storing job ids returned from scrapyd in DB. My spiders are storing items in DB.

Question is: how could I link in DB the job id issued by scrapyd and items issued by the crawl?

I could trigger my spider using an extra parameter - let say an ID generated by my web application - but I'm not sure it is the best solution. At the end, there is no need to create that ID if scrapyd issues it already...

Thanks for your help

2 Answers

In the spider constructor(inside init), add the line -->

self.jobId = kwargs.get('_job')

then in the parse function pass this in item,

def parse(self, response):
    data = {}
    ......
    yield data['_job']

in the pipeline add this -->

def process_item(self, item, spider):
    self.jobId = item['jobId']
    .......
Related