How to get dynamic html with Selenium?

Viewed 344

I'm trying to develop a web crawler in python that, given a website, analyzes its html and searches for all href tags, but with libraries like Beautiful Soap it is not possible to obtain the dynamic content of an html page, in fact the crawler I am making must also discover the hrefs generated by any script for example. So I discovered Selenium and made this script:

driver = webdriver.Chrome()
driver.get(url)
driver.execute_script("return document.body.innerHTML")
time.sleep(15)
html = driver.page_source
print("HTML :", html)
links = []
elements = driver.find_elements_by_tag_name('a')
for elem in elements:
    href = elem.get_attribute("href")
    links.append(href)
return links

But when I run it I don't find in the html the content that, for example with Chrome developer tools, I see, so my question is: how can I get the whole html of a page with also the html generated by a generic script?

URL TO TEST: "https://www.lubecreostorepratolapeligna.it/it/cucine-lube/cucine-moderne/"

Example of test: I want to take the href of an image of a kitchen present in the catalog

N.B. I don't want to select an element and wait for it thanks to WebDriverWait as I'm creating a generic crawler for any site so I don't have a particular element to wait or search for, I just want to get the dynamic content of a generic html.

If there are better libraries for my purpose please tell me.

UPDATE : i have found a solution to my problem here is the code to search for any iframe in the html (dynamic content of page) and then navigate them

options = Options()
options.add_argument('--headless')
browser = webdriver.Chrome(options=options)
browser.get(url_to_search_for)
soup = BeautifulSoup(browser.page_source, "html.parser")
browser.close()
        
iframe = []
for x in soup.find_all('iframe'):
    print(x['src'])
    if str in x['src']:
        print('ciao')
        iframe.append(x['src'])
for x in iframe:
    try:
        page = urllib.request.urlopen(x, timeout=20)        
    except HTTPError as e:
        page = e.read()
            
    soup = BeautifulSoup(page, 'html.parser')
    for a in soup.find_all('a', href=True):
        print("HREF IFRAME", a['href'])
1 Answers

The elements you are looking after is present inside an iframe.You need to switch to iframe first in order to get those elements.

url="https://www.lubecreostorepratolapeligna.it/it/cucine-lube/cucine-moderne/"
driver = webdriver.Chrome()
driver.get(url)
time.sleep(5)
driver.switch_to.frame(driver.find_element_by_css_selector("div.row iframe"))
time.sleep(5)
elements=driver.find_elements_by_css_selector(".ajax_content div.gb_CL_product>a")
links=[]
for elem in elements:
    href = elem.get_attribute("href")
    links.append(href)

print(links)

Please Note:- time.sleep() is a bad practice you should use WebDriverWait()

Console Output:

['https://www.cucinelube.it/it/cucine-moderne/adele-project/', 'https://www.cucinelube.it/it/cucine-moderne/clover/', 'https://www.cucinelube.it/it/cucine-moderne/creativa/', 'https://www.cucinelube.it/it/cucine-moderne/essenza/', 'https://www.cucinelube.it/it/cucine-moderne/gallery/', 'https://www.cucinelube.it/it/cucine-moderne/georgia/', 'https://www.cucinelube.it/it/cucine-moderne/immagina-plus/', 'https://www.cucinelube.it/it/cucine-moderne/luna/', 'https://www.cucinelube.it/it/cucine-moderne/noemi/', 'https://www.cucinelube.it/it/cucine-moderne/oltre/', 'https://www.cucinelube.it/it/cucine-moderne/round/', 'https://www.cucinelube.it/it/cucine-moderne/swing/']
Related