How can I scrape remote jobs on Stack Overflow Jobs with Scrapy?

Viewed 212

Dear fellow software engineers,

I have problems with scraping remote jobs on Stack Overflow Jobs. I have googled for a solution and have studied parts of the Scrapy documentation, but I haven't found the solution to my problems.

My goal is to scrape remote jobs on Stack Overflow Jobs.

Here are the fields that I want to get for each remote job:

  • Job title
  • Salary range (if present)
  • Job type
  • Experience level
  • Role
  • Industry (if present)
  • Company size (if present)
  • Company type (if present)
  • Preferred timezone (if present)
  • Technologies

I have tried doing this with 2 different spiders. Below is each one of them with brief explanation:

The first spider I wrote mimicked the URL structure I have in my browser as I browse remote jobs. For example, if I'm on the third page of the remote jobs listings, I get the URL that looks like the following: https://stackoverflow.com/jobs?id=522484&pg=3&r=true I mimicked that URL structure in my spider.

Here's the code for that spider:

import scrapy
import time

class JobsSpider(scrapy.Spider):
    name = "jobs"
    start_urls = [
        "https://stackoverflow.com/jobs/remote-developer-jobs"
    ]
    already_visited_links = []

    def parse(self, response, page_number=0): # 0 is a magic number for an invalid page number
        jobs = response.xpath("//div[contains(@class, 'job')]")
        links_to_next_pages = response.xpath("//a[contains(@class, 's-pagination--item')]").css("a::attr(href)").getall()

        # visit each job page (as I do in the browser) and scrape the relevant information (Job title etc.)
        for job in jobs:
            # get the job ID for each job
            job_id = int(job.xpath('@data-jobid').extract_first()) # there will always be one element
            # now visit the link with the job_id and get the info
            if page_number != 0:
                job_link_to_visit = "https://stackoverflow.com/jobs?id=" + str(job_id) + "&r=true" + "&p=" + str(page_number)
            else:
                job_link_to_visit = "https://stackoverflow.com/jobs?id=" + str(job_id) + "&r=true"
            request = scrapy.Request(job_link_to_visit,
                             callback=self.parse_job)
            yield request

        # go to the next job listings page (if you haven't already been there)
        # not sure if this solution is the best since it has a loop which has a recursion in it
        for link_to_next_page in links_to_next_pages:
            if link_to_next_page not in self.already_visited_links:
                self.already_visited_links.append(link_to_next_page)
                next_page_number = 0
                # below is a pattern I noticed: whenever a link has a p in its GET request, it is always near the end
                # therefore, here I check if the last thing in the link is "p=<someNumber>" and if it is
                # then I get <someNumber>; I know this is not robust, but I currently can't think of a better way to do this
                if link_to_next_page[-4] == "p":
                    next_page_number = int(link_to_next_page[-1])
                if link_to_next_page[-5] == "p":
                    next_page_number = int(link_to_next_page[-2:])
                yield response.follow(link_to_next_page, callback=self.parse, cb_kwargs={"page_number" : next_page_number})

        print("End of parse method")

    def parse_job(self, response):
        with open("dump.txt", "wb") as f:
            f.write(response.body)

However, the parse_job method of the above code produced a dump.txt file which doesn't match the contents of the website I saw in my web browser.

The approach for my 2nd spider was to visit the job details page directly, i.e. https://stackoverflow.com/jobs/522636, but that didn't work as well. The contents of the webpage being scraped in my web browser and the contents of the dump.txt file differ. Here's the code for that spider:

import scrapy
import time

class JobsSpider(scrapy.Spider):
    name = "jobs"
    start_urls = [
        "https://stackoverflow.com/jobs/remote-developer-jobs"
    ]
    already_visited_links = []

    def parse(self, response):
        jobs = response.xpath("//div[contains(@class, 'job')]")
        links_to_next_pages = response.xpath("//a[contains(@class, 's-pagination--item')]").css("a::attr(href)").getall()

        # visit each job page (as I do in the browser) and scrape the relevant information (Job title etc.)
        for job in jobs:
            # get the job ID for each job
            job_id = int(job.xpath('@data-jobid').extract_first()) # there will always be one element
            # now visit the link with the job_id and get the info
            job_link_to_visit = "https://stackoverflow.com/jobs/" + str(job_id)
            request = scrapy.Request(job_link_to_visit,
                             callback=self.parse_job)
            yield request

        # go to the next job listings page (if you haven't already been there)
        # not sure if this solution is the best since it has a loop which has a recursion in it
        for link_to_next_page in links_to_next_pages:
            if link_to_next_page not in self.already_visited_links:
                self.already_visited_links.append(link_to_next_page)
                yield response.follow(link_to_next_page, callback=self.parse)

        print("End of parse method")

    def parse_job(self, response):
        with open("dump.txt", "wb") as f:
            f.write(response.body)

From my understanding, I think that my problem stems from the fact that there's some JavaScript and/or AJAX here that somehow makes the server know that my spider is indeed a spider. Again, I have looked around, but failed to find the solution to my problems.

Can someone tell me how can I build a spider that scrapes the information I want to scrape from remote jobs on Stack Overflow Jobs?

Thank you!

0 Answers
Related