Why can't I get content of the full page using requests_html with Python?

Viewed 191

I was trying to get data from this website https://aptekazawiszy.pl/ with Python. I have to search for a product using their search engine and if the product exists, extract direct link to it, make a request to that site and then get some data of it. The problem is, I don't receive full HTML of the search results page. I was using urllib3 requests and BeautifulSoup4. I red that the site could be rendered with JavaScript and I'll have to render it by myself so I used requests_html library which supports JS. Unfortunetly, it didn't work either. The example code below I wrote should print the first product of this site https://aptekazawiszy.pl/strony/wyszukiwanie.html?name=myd%C5%82o%20z%20olejem%20konopnym&pf-size=32&pf-page=1, but prints None.

from requests_html import HTMLSession

session = HTMLSession()

def get_prod_link(prod):
    p = prod[0].replace(' ','%20')
    url = f'https://aptekazawiszy.pl/strony/wyszukiwanie.html?name={p}&pf-size=32&pf-page=1'
    r = session.get(url)
    r.html.render(sleep=1, keep_page=True, scrolldown=1)
    xpath = '//*[@id="content"]/aside/div[10]/div[2]/div[1]/article/div/figure/a/img'
    elem = r.html.xpath(xpath) 
    return elem


if __name__ == '__main__':
    product = "mydlo z olejem konopnym"
    print(get_prod_link(product))
0 Answers
Related