Scraping Reelgood.com in python with Beautifullsoup

Viewed 206

Im trying to build a scraper (in Python) for the website ReelGood.com.

If I go to a specific movie on Reelgood it shows me a play button like this: Stream button

If I click that button it redirects me for example to https://www.netflix.com/title/70232180 now I would like to scrape that specific URL. so I thought I make a small python script to scrape all links that contains https://netflix.com/x.

so I came up with this:

from bs4 import BeautifulSoup
import requests

URL = "https://reelgood.com/movie/the-intouchables-2011"
page = requests.get(URL)
soup = BeautifulSoup(page.content, "html.parser")
for a_href in soup.find_all("a", href=True):
    print(a_href["href"])

now this does give me a print from all the links but there are no links containing the url witch im redirected to.

anyone got an idea how to spit out the Netflix.com url?

2 Answers

To filter for links that only contain https://www.netflix.com/, you can use a CSS Selector: a[href*="https://www.netflix.com/"], which will select all a with an href containing https://www.netflix.com/.

from bs4 import BeautifulSoup
import requests

URL = "https://reelgood.com/movie/the-intouchables-2011"
page = requests.get(URL)
soup = BeautifulSoup(page.content, "html.parser")

for a_href in soup.select('a[href*="https://www.netflix.com/"]'):
    print(a_href["href"])

Output:

https://www.netflix.com/watch/70232180

I think that you can use my code addapt it to your needs
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import WebDriverWait import time from geopy.geocoders import Nominatim import time from pprint import pprint

# instantiate a new Nominatim client
app = Nominatim(user_agent="tutorial")

def getLocation():
    #autoriser le naviagateur pour acceder à l'emplacement actuelle par defaut,
    # Si on essaye d'accéder à un site Web : « https://mycurrentlocation.net » via chrome,
    # il demande d'autoriser l'accès à la localisation. La commande « - use-fake-ui-for-media-stream »
    # accordera toutes les autorisations pour l'emplacement, le microphone, etc. automatiquement.
    options = Options()
    options.add_argument("--use-fake-ui-for-media-stream")
    # appelez la page Web https://mycurrentlocation.net/ et attendez 20 secondes que la page se charge.
    timeout = 20
    #Pour chromedriver il faut avoir la meme version que google chrome
    driver = webdriver.Chrome(executable_path = './chromedriver.exe', chrome_options=options)
    #driver = webdriver.Chrome(executable_path = './chromedriver.exe', chrome_options=options) -> on peut mettre ça a la palce
    driver.get("https://mycurrentlocation.net/")
    wait = WebDriverWait(driver, timeout)
    time.sleep(3)
    #Trouvez le XPath des éléments de latitude et de longitude mentionnés sur la page Web puis copier le nom de la classe qu'on souhaite récupérer
    neighborhood = driver.find_elements_by_xpath('//*[@id="neighborhood"]')
    neighborhood = [x.text for x in neighborhood]
    neighborhood = str(neighborhood[0])

    regionname = driver.find_elements_by_xpath('//*[@id="regionname"]')
    regionname = [x.text for x in regionname]
    regionname = str(regionname[0])

    placename = driver.find_elements_by_xpath('//*[@id="placename"]')
    placename = [x.text for x in placename]
    placename = str(placename[0])

    driver.quit()
    return (neighborhood,regionname,placename)

neighborhood,regionname,placename=getLocation()

print("le résultat est : \n ",
    neighborhood,regionname,placename)
Related