I have been trying to download a PDF file using requests but, no matter what I do, it keeps returning 403 as status and it is impossible to open the downloaded PDF.
Here is the code I am running:
import requests
url_pdf='https://www.agerborsamerci.it/wp-content/uploads/2022/01/Settimanale-n.-2-del-20-Gennaio-2022-%E2%80%93-Listino-Borsa-n.-2.pdf'
#session = requests.Session()
headers = {
"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/97.0.4692.99 Safari/537.36",
"Accept": "image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8",
"Cache-Control": "image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8",
"host-header": "6b7412fb82ca5edfd0917e3957f05d89",
"Accept-Encoding": "gzip, deflate, br",
"cache-control": "image/avif,image/webp,image/apng,image/svg+xml,image/*,*/*;q=0.8",
"Connection": "keep-alive",
"referer":"https://www.agerborsamerci.it/wp-content/uploads/2022/01/Settimanale-n.-2-del-20-Gennaio-2022-%E2%80%93-Listino-Borsa-n.-2.pdf"
}
req=requests.get(url_pdf, headers=headers)
print(req.status_code)
with open("bologna.pdf", 'wb') as f:
f.write(req.content)
f.closed
As you can see, I have tried using a 'Session' object, setting (different) 'User-Agent' as well as other headers but nothing seems to work.
I have also tried using
import os
name='bologna.pdf'
os.system('wget {} -O {}'.format(url_pdf,name))
But it is not working either.
Do you have any idea about what could I do to overcome this problem? I am really struggling to figure it out.
Thank you a lot!
