Scrapy/Selenium: How do I follow the links in 1 webpage?

Viewed 133

I am new to web-scraping.

  1. I want to go to Webpage_A and follow all the links there.
  2. Each of the link lead to a page where I can select some button and download the data in an Excel file.

I tried the below code. But I believe there is an error with

            if link:
                yield SeleniumRequest(

Instead of using "SeleniumRequest" to follow the links, what should I use?
If using pure Scrapy, I know I can use

yield response.follow(

Thank you

class testSpider(scrapy.Spider):
    name = 'test_s'
    
    def start_requests(self):
        yield SeleniumRequest(
            url='CONFIDENTIAL',
            wait_time=15,
            screenshot=True,
            callback=self.parse
        )

    def parse(self, response):
       
        tables_name = response.xpath("//div[@class='contain wrap:l']//li")
        for t in tables_name:
            name=t.xpath(".//a/span/text()").get()
            link = t.xpath(".//a/@href").get()
   
            if link:
                yield SeleniumRequest(
                    meta={'table_name': name},
                    url= link,
                    wait_time=15,
                    screenshot=True,
                    callback=self.parse_table
                )
            
                

    def parse_table(self, response):
     
        name = response.request.meta['table_name']
        button_select=response.find_element_by_xpath("(//a[text()='Select All'])").click()       
        button_st_yr=response.find_element_by_xpath("//select[@name='ctl00$ContentPlaceHolder1$StartYearDropDownList'] /option[1]").click()
        button_end_mth=response.find_element_by_xpath("//select[@name='ctl00$ContentPlaceHolder1$EndMonthDropDownList']/option[text()='Dec']").click()
        
        button_download=response.find_element_by_xpath("//input[@id='ctl00_ContentPlaceHolder1_DownloadButton']").click() 
        
        yield{
            'table_name':  name 
            }

0 Answers
Related