I am trying to scrape a dynamic page using BeautifulSoup. After accessing the said page from https://www.nemlig.com/ with the help of Selenium (and thanks to the code advice from @cruisepandey) like this:
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
driver = webdriver.Chrome(executable_path = r'C:\Users\user\lib\chromedriver_77.0.3865.40.exe')
wait = WebDriverWait(driver,10)
driver.maximize_window()
driver.get("https://www.nemlig.com/")
wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, ".timeslot-prompt.initial-animation-done")))
wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, "input[type='tel'][class^='pro']"))).send_keys('2300')
wait.until(EC.element_to_be_clickable((By.CSS_SELECTOR, ".btn.prompt__button"))).click()
I am prompted with this page that I want to scrape.
More precisely, at this point, I want to scrape the rows from the right-hand side of the page. If you look through the HTML code behind these you will notice that the div class time-block__row has 3 different data-automation attributes for the main 3 times of the day.
<div class="time-block__row" data-automation="beforDinnerRowTmSlt">
<div class="time-block__row-header">Formiddag</div>
<div class="no-timeslots ng-hide" ng-show="$ctrl.timeslotDays[$ctrl.selectedDateIndex].morningHours == 0">
Ingen levering..
</div>
<!----><!----><div class="time-block__item duration-1 disabled" ng-repeat="item in $ctrl.selectedHours track by $index" ng-if="item.StartHour >= 0 && item.StartHour < 12" ng-click="$ctrl.setActiveTimeslot(item, $index)" ng-class="['duration-1', {'cheapest': item.IsCheapHour, 'event': item.IsEventSlot, 'selected': $ctrl.selectedTimeId == item.Id || $ctrl.selectedTimeIndex == $index, 'disabled': item.isUnavailable()}]" data-automation="notActiveSltTmSlt">
<div class="time-block__inner-container">
<div class="time-block__time">8-9</div>
<div class="time-block__attributes">
<!----></div>
<div class="time-block__cost">29 kr.</div>
So Formiddag (Morning) has data-automation = "beforDinnerRowTmSlt", Eftermiddag (Afternoon) has data-automation = "afternoonRowTmSlt" and Aften (Evening) has data-automation = "eveningRowTmSlt".
page_source = wait.until(driver.page_source)
soup = BeautifulSoup(page_source)
time_of_the_day = soup.find('div', class_='time-block__row').text
- The problem is
using the code above, time_of_the_day only contains information from the Morning rows.
How can I scrape these rows properly using the data-automation attribute? How can I possibly access the other 2 div classes and their child divs? My plan is to create a dataframe containing something like this:
Time_of_the_day Hours Price Day
Formiddag 8-9 29kr. Tor. 10/10
.... .... .... ....
Eftermiddag 12-13 29kr. Tor. 10/10
.... .... .... ....
The day column will contain the output from here: day = soup.find('div', class_='content').text
I know this is quite a lengthy post but hopefully I've made it easy to understand the task and you will be able to help me out with advice, tips or code!
