I am working on a webscraper using html requests and beautiful soup (I am new to this). For 1 webpage (https://www.selfridges.com/GB/en/cat/beauty/make-up/?pn=1) I am trying to scrape the links of each product in a product grid. I have tried using absolute_links and the xpath:
session = HTMLSession()
for x in range(1, 30):
url = f'https://www.selfridges.com/GB/en/cat/beauty/make-up/?pn={x}'
r = session.get(url)
r.html.render(sleep=2)
products = r.html.xpath('//*[@id="content"]/div[3]/div/div/div/div[2]/div[1]/div[2]/div[6]/div/div/div[1]/div/div/div/div[2]', first=True)
productlist = products.absolute_links
productlinks.extend(productlist)
print(productlinks)
and BeautifulSoup:
session = HTMLSession()
for x in range(1, 30):
url = f'https://www.selfridges.com/GB/en/cat/beauty/make-up/?pn={x}'
r = session.get(url)
r.html.render(sleep=2)
soup = BeautifulSoup(r.content, 'lxml')
productlist = r.html.find('div', class_="listing-items c-listing-items initialized")
print(productlist)
for item in productlist:
for link in item.find_all('a', href=True):
productlinks.append(baseurl + link['href'])
print(productlinks)
Both return Empty lists or an AttributeError: 'NoneType' object has no attribute 'absolute_links'. I am unsure of why this happens. Any help would be appreciated.