How to retreive till the last page of the table?

Viewed 41

In the code below I don't know where is the last page, the code below works till PAGE 25 which I mentioned manually! sometimes we have 60 or 70 pages! How can I change the code and get the table till the last page??

   from selenium import webdriver
import pandas as pd
import time
driver = webdriver.Chrome('C:\Webdriver\chromedriver.exe')
driver.get('https://www150.statcan.gc.ca/n1/pub/71-607-x/2021004/imp-eng.htm?r1=(1)&r2=0&r3=0&r4=12&r5=0&r7=0&r8=2022-01-01&r9=2022-01-01')
time.sleep(2)

Canada_Result=[]
for J in range (25):
    commodities = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[2]/a')
    Countries = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[4]')
    quantities = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[7]')
    weights = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[8]/abbr')

    
    for i in range(25):
        temporary_data= {'Commodity': commodities[i].text,'Country': Countries[i].text,'quantity': quantities[i].text, 'weight': weights[i].text }
        Canada_Result.append(temporary_data)
    df_data = pd.DataFrame(Canada_Result)
    df_data
    df_data.to_excel('Canada_scrapping_result.xlsx', index=False)
    # click on the Next button    
    driver.find_element_by_xpath('//*[@id="report_results_next"]').click()
    time.sleep(1)
1 Answers

I used the variable pages to find how many buttons there were and then work out how many pages there are by getting the text from the last non "Next" button on the page. This is great for dealing with multi-page selections but I have also implemented a solution in case the selection only has one page e.g. less than 25 rows retrieved.

I added an if statement near the start for the categories that have less rows than 25 for example this one https://www150.statcan.gc.ca/n1/pub/71-607-x/2021004/imp-eng.htm?r1=(1)&r2=9&r3=1&r4=02&r5=0&r7=0&r8=2022-01-01&r9=2022-04-01 which only has 19 rows retrieved.

The variable period_entries is just used to see how many row entries there are on each page. I implemented this mainly due to the last page having only 22 entries instead of 25 which broke the program initially when it got to the end.

The last if statement is there to ensure the program still scrapes the very last page but does not try to click the next button since it is not available.

from selenium import webdriver
import pandas as pd
import time

driver = webdriver.Chrome()
driver.get('https://www150.statcan.gc.ca/n1/pub/71-607-x/2021004/imp-eng.htm?r1=(1)&r2=0&r3=0&r4=12&r5=0&r7=0&r8=2022-01-01&r9=2022-01-01')
time.sleep(2)

Canada_Result=[]
pages = len(driver.find_elements_by_xpath('//a[@class="paginate_button" or @class="paginate_button current" and @title]'))
pages = driver.find_element_by_xpath('//a[@onclick and @class="paginate_button" or @class="paginate_button current" and @title][%d]' % (pages)).text.strip("Page\n")

if pages == '':
    pages = 1

for J in range (int(pages)):
    commodities = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[2]/a')
    Countries = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[4]')
    quantities = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[7]')
    weights = driver.find_elements_by_xpath('.//*[@id="report_table"]/tbody/tr["i"]/td[8]/abbr')

    period_entries = len(commodities)
    
    for i in range(period_entries):
        temporary_data= {'Commodity': commodities[i].text,'Country': Countries[i].text,'quantity': quantities[i].text, 'weight': weights[i].text }
        Canada_Result.append(temporary_data)
    df_data = pd.DataFrame(Canada_Result)
    df_data
    df_data.to_excel('Canada_scrapping_result.xlsx', index=False)

    if J == int(pages) - 1:
        print("Done")
        break
    # click on the Next button
    driver.find_element_by_xpath('//*[@id="report_results_next"]').click()
    time.sleep(1)
Related