get link after href-Tag with BeautifulSoup

Viewed 83

i want to download all the pictures from this side in high resolution and not the preview pictures:

https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl

The link -> https://xxxxx.de to the images i want to download is stored in this part of the html-code: link to the picture

The Code i tried so far was that:

from bs4 import BeautifulSoup
import requests
page = requests.get("https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl")

souped = BeautifulSoup(page.content, "html.parser")
for pic in souped.find_all(class_="preview hasXXL"):   
   print(pic['href'])

With that i get to the right part of the code. But i don't get it how to scrape the link after the href-tag. When i want to scarpe it i get that results:

/app/detail.php?id=A02A8f9001ZZl&picNo=1" id="preview_1

But i expect that:

https://images.booklooker.de/x/02Sh07/Donna-W-Cross+Die-P%C3%A4pstin.jpg

What did i do wrong?

Thanks a lot for your help!!

3 Answers

If you want the image URLs (e.g. https://images.booklooker.de/t/02Sh07/Donna-W-Cross+Die-P%C3%A4pstin.jpg) then you'd need to follow the previewImage elements in the HTML (not the "preview hasXXL" class) and extract the "src" attribute from the img element for the URL.

from bs4 import BeautifulSoup
import requests

url = "https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl"
page = requests.get(url)

souped = BeautifulSoup(page.content, "html.parser")

for pic in souped.find_all("img", class_="previewImage"):   
    src = pic['src']

    # Next change thumbnail URL to full res image
    src = src.replace("/t/", "/s/")

    print(src)

Output:

https://images.booklooker.de/s/02Sh07/Donna-W-Cross+Die-P%C3%A4pstin.jpg
...
https://images.booklooker.de/s/02Sh0S/Donna-W-Cross+Die-P%C3%A4pstin.jpg

Since you're trying to download images, you may search for the <img> tag and utilise it's src attribute which provides the accurate information.

Your Modified Code:

Method 1: Searching for <img> with class="prevewImage" after reviewing the html code.

from bs4 import BeautifulSoup
import requests

page = requests.get("https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl")

souped = BeautifulSoup(page.content, "html.parser")
for pic in souped.find_all("img", class_="previewImage"): # <img src>
    print(pic["src"])

Method: Using the .img to find the image tag that is associated with the searching <a>.

from bs4 import BeautifulSoup
import requests

page = requests.get("https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl")

souped = BeautifulSoup(page.content, "html.parser")
for pic in souped.find_all(class_="preview hasXXL"): # <a> <img src> </a>
    print(pic.img["src"])

Output:

https://images.booklooker.de/t/02Sh07/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/t/02Sh08/Donna-W-Cross+Die-P%C3%A4pstin.jpg
...
https://images.booklooker.de/t/02Sh0S/Donna-W-Cross+Die-P%C3%A4pstin.jpg

This is how you get the max res images from there:

from bs4 import BeautifulSoup
import requests
import json

url = 'https://www.booklooker.de/B%C3%BCcher/Donna-W-Cross+Die-P%C3%A4pstin/id/A02A8f9001ZZl'
r = requests.get(url)
soup = BeautifulSoup(r.text, 'html.parser')
data_obj = json.loads(soup.select_one('div#imageWrapper').get('data-imageinfo'))['images']
for x in data_obj:
    print(data_obj[x]['url_xxl'])

Result:

https://images.booklooker.de/x/02Sh07/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh08/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh09/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh0C/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh0F/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh0M/Donna-W-Cross+Die-P%C3%A4pstin.jpg
https://images.booklooker.de/x/02Sh0S/Donna-W-Cross+Die-P%C3%A4pstin.jpg
Related