I have a script that visits the website of a local weather station every hour and scrapes the current rainfall. This gets fed in my sprinkler server/database/etc. Sometimes there is a hiccup in the network connection or something on the side of the website. such an unfortunate makes the method hang indefinitely. This is not good for the stability of my program.
I've experimented with multiprocessing but I couldn't get it to work properly. Ideally it would launch the scraping module, and returns every second for a maximum of 10 so see if the method has stopped and produced an output. If the 10 sec is exceeded it should kill it and try the next hour. How would you tackle this issue?
My current script is this:
from multiprocessing import Process,Queue,Pipe
import time
import requests
import urllib.request
from bs4 import BeautifulSoup
url = "https://www.weerstationzoersel.be/weather2/index.php?p=10"
def get_rain():
try:
response = requests.get(url)
responsestr = str(response)
if "200" in responsestr:
soup = BeautifulSoup(response.text, "html.parser")
tags = soup.findAll('span')
line_rain = str(tags[15])
line_rain = line_rain[62::]
rainfall = line_rain.rstrip("</span>")
rainfall = round(float(rainfall.replace(',','.')),1)
except:
rainfall="error"
return(rainfall)
if __name__ == '__main__':
print(get_rain())