How to scrape the job postings from https://apply.workable.com/caxton using Selenium with Python?

Viewed 66

I am trying to scrape job postings from this website: https://apply.workable.com/caxton and am using Selenium with Python for this exercise.

Here is the portion of the website's HTML where I face the problem:

enter image description here

I am trying to reach the <li> tags inside the <main> tag but get the Unable to locate element

error when I try to find the <main> tag using XPATH, TAG NAME, CLASS NAME or CSS SELECTOR. Please see my code and the result below.

Code

Output

Things seem to be fine until //div[@id = 'app']//div//div, since printing elem1 gives the webdriver element as the output (please see below).

Code

Output

Could you please suggest what code I can use to reach the <main> tag and ultimately the <li> tags contained therein?

3 Answers

This would give you the main

driver.find_element(By.XPATH, "//*[@role='main']")

DOM Snapshot

If you are looking for the job opening cards, this may help:

driver.find_elements(By.XPATH, "//*[@role='main']//a")

OR

driver.find_elements(By.XPATH, "//*[@role='main']//li[@data-ui = 'job-opening']")

UPDATE TO SHOW THE COMPLETE CODE:

To get only the main element:

driver.get("https://apply.workable.com/caxton")
main_ele = WebDriverWait(driver, 30).until(EC.visibility_of_element_located((By.XPATH, "//*[@role='main']")))
print(main_ele)

Output:

<selenium.webdriver.remote.webelement.WebElement (session="34757b3f076c7ad292e832b683654e29", element="73ee8fcc-807d-4066-a884-46f172d859bb")>

Process finished with exit code 0

Note that main does not contain any text to render and hence webelement is seen in output.

To get the job cards:

driver.get("https://apply.workable.com/caxton")
job_cards = WebDriverWait(driver, 30).until(EC.visibility_of_all_elements_located((By.XPATH, "//*[@role='main']//li[@data-ui = 'job-opening']")))
jobs = [job.text for job in job_cards]
print(jobs)

Output:

['Posted 4 days ago\nEmerging Markets Macro Portfolio Manager\nNew York, New York, United States', 'Posted 4 days ago\nEquity L/S Portfolio Manager\nNew York, New York, United States', 'Posted 4 days ago\nGlobal Macro Portfolio Manager\nNew York, New York, United States', 'Posted 4 days ago\nFixed Income RV Portfolio Manager\nNew York, New York, United States', 'Posted 4 days ago\nFixed Income RV Portfolio Manager\nLondon, England, United Kingdom', 'Posted 13 days ago\nESG Internship\nLondon, England, United Kingdom', 'Posted 26 days ago\nTreasury Operations Supervisor\nLondon, England, United Kingdom', 'Posted about 1 month ago\nCorporate Receptionist & Office Management Assistant\nLondon, England, United KingdomFull time', 'Posted about 2 months ago\nFund Accounting, Vice President\nLondon, England, United KingdomFull time', 'Posted 3 months ago\nFund Accounting, Associate\nNew York, New York, United StatesFull time']

Process finished with exit code 0

Here, there are about 10 job-cards in the page, and first, all the elements are located using visibilit_of_all_elements_located and then looped through each of them to extract the text and appended into a list, which is output finally.

To handle dynamic elements you need to use WebDriverWait() and wait for visibility of all elements.

Use following css selector to identify the li element.

driver.get("https://apply.workable.com/caxton/")
elements=WebDriverWait(driver,10).until(EC.visibility_of_all_elements_located((By.CSS_SELECTOR, "[role='main']  li[data-ui='job-opening']")))
print("Total li elements on the page: " + str(len(elements)) )
for ele in elements:
    print(ele.text)

you need to import below libraries.

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

output on my terminal:

enter image description here

Within the website you can identify all the Job Postings texts using the following locator strategy:

  • xpath:

    //li[@data-ui='job-opening']//ancestor::div[1]//h3
    

Hence, you can using list comprehension and your line of code will be:

print([my_elem.text for my_elem in driver.find_element(By.XPATH, "//li[@data-ui='job-opening']//ancestor::div[1]//h3")])
  

Ideally you need to induce WebDriverWait for the visibility_of_all_elements_located() and you can use:

print([my_elem.get_attribute("innerHTML") for my_elem in WebDriverWait(driver, 20).until(EC.visibility_of_all_elements_located((By.XPATH, "//li[@data-ui='job-opening']//ancestor::div[1]//h3")))])

Console Output:

['Emerging Markets Macro Portfolio Manager', 'Equity L/S Portfolio Manager', 'Global Macro Portfolio Manager', 'Fixed Income RV Portfolio Manager', 'Fixed Income RV Portfolio Manager', 'ESG Internship', 'Treasury Operations Supervisor', 'Corporate Receptionist & Office Management Assistant', 'Fund Accounting, Vice President', 'Fund Accounting, Associate']

Note : You have to add the following imports :

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
Related