Why my scrapper is scrapping only 5 images then returning errors

Viewed 57

I coded this web scraper, and it is supposed to download all the images from this url: https://www.olx.ro/d/oferte/q-iphone-13/ however it downloads only five images then it returns errors for the rest of them like this : enter image description here

Here is my code:

import bs4
import requests
import urllib.request
from urllib.request import urlopen
from bs4 import BeautifulSoup


url="https://www.olx.ro/d/oferte/q-iphone-13/"

page=urllib.request.urlopen(url)

page_soup=BeautifulSoup(page,'html.parser')
test=page_soup.find_all('div', class_="css-19ucd76")

i=1

for img in test:
    try:
        img_tag=img.find('img')
        img_src=img_tag.get('src')
        image=img_src
        if(image!='/app/static/media/no_thumbnail.15f456ec5.svg'):

            print(image)
        else:
            print('error')
        file_name=str(i)
        i+=1
        ''' img_file=open(file_name +'.jpeg','wb')
        img_file.write(urllib.request.urlopen(image).read())
        img_file.close()'''
    except :
        
        pass

please, help. Thank you!

1 Answers

Try to add headers like this

import requests
from bs4 import BeautifulSoup


headers = {
    'Access-Control-Allow-Origin': '*',
    'Access-Control-Allow-Methods': 'GET',
    'Access-Control-Allow-Headers': 'Content-Type',
    'Access-Control-Max-Age': '3600',
    'User-Agent': 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:52.0) Gecko/20100101 Firefox/52.0'
    }

url = "https://example.com"
req = requests.get(url, headers)
soup = BeautifulSoup(req.content, 'html.parser')
print(soup.prettify())

Related