Web scraping using Beautifulsoup/Selenium unable to pull div class using find or find_all

Viewed 208

I am trying to webscrape form the website https://roll20.net/compendium/dnd5e/Monsters%20List#content and having some issues.

My first script I tried kept returning an empty list when finding by div and class name, which I believe is do to the site using Javascript? But a little uncertain if that is the case or not.

Here was my first attempt:

import requests
from bs4 import BeautifulSoup
import pandas as pd

page = requests.get('https://roll20.net/compendium/dnd5e/Monsters%20List#content')

soup = BeautifulSoup(page.text, 'html.parser')

card = soup.find_all("div", class_='card')

print(card)

This one returns an empty list so then I tired to use Selenium and scrape with that. Here is that script:

import requests
from bs4 import BeautifulSoup
import pandas as pd
from selenium import webdriver

url='https://roll20.net/compendium/dnd5e/Monsters%20List#content'
driver = webdriver.Firefox(executable_path = 'C:\Windows\System32\geckodriver')
driver.get(url)

page = driver.page_source
page_soup = soup(page,'html.parser')

Starting the script with that I then tried all 3 of these different options (individually ran these, just listed them here together for simplicity sake):

for card in body.find('div', {"class":"card"}):
    print(card.text)

print(card)
    
for card in body.find_all('div', {"class":"card"}):
    print(card.text)

print(card)        

card = body.find_all('div', {"class":"card"})

print(card)

All of them return the same error message: AttributeError: ResultSet object has no attribute 'find'. You're probably treating a list of elements like a single element. Did you call find_all() when you meant to call find()?

Where am I going wrong here?

Edit:

Fazul thank you for your input on this I guess I should be more specific. I was more looking to get the contents of each card. For example, the card has a "body" class and within that body class there are many fields that is the data I am looking to extract. Maybe I am misunderstanding your script and what you stated. Here is a screen shot to maybe help specify my question a bit more to what content I am looking to extract.

enter image description here

So everything that would be under the body i.e. name, title, subtitle, etc.. Those were the texts I was trying to extract.

3 Answers

That page is being loaded by JavaScript. So beautifulsoup will not work in this case. You have to use Selenium.

And the element that you are looking for - <div> with class name as card show up only when you click on the drop-down arrow. Since you are not doing any click event in your code, you get an empty result.

  • Use selenium to click the <div> with class name as dropdown-toggle. That click event loads the <div class='card'>
  • Then you can scrape the data you need.

Since the page is loaded by JavaScript, you have to use Selenium.

My solution As follows:

Every card has a link as shown: imge

Use BeautifulSoup to get the link for each card, Then open each link using Selenium in headless mode because it is also loaded by JavaScript.

Then you get the data you need for each card using BeautifulSoup

Working code:

from selenium import webdriver
from webdriver_manager import driver
from webdriver_manager.chrome import ChromeDriver, ChromeDriverManager
from selenium.webdriver.chrome.options import Options
from time import sleep
from bs4 import BeautifulSoup

driver = webdriver.Chrome(ChromeDriverManager().install())
driver.get("https://roll20.net/compendium/dnd5e/Monsters%20List#content")
page = driver.page_source

# Wait until the page is loaded
sleep(13)

soup = BeautifulSoup(page, "html.parser")

# Get all card urls
urls = soup.find_all("div", {"class":"header inapp-hide single-hide"})

cards_list = []
for url in urls:
    card = {}
    card_url = url.a["href"]
    card_link = f'https://roll20.net{card_url}#content'

    # Open chrom in headless mode
    chrome_options = Options()
    chrome_options.add_argument("--headless")
    driver = webdriver.Chrome(ChromeDriverManager().install(), options=chrome_options)
    driver.get(card_link )
    page = driver.page_source
    soup = BeautifulSoup(page, "html.parser")
    card["name"] = soup.select_one(".name").text.split("\n")[1]
    card["subtitle"] = soup.select_one(".subtitle").text
    # Do the same to get all card detealis 

    cards_list.append(card)
    driver.quit()
   

Library you need to install

pip install webdriver_manager

This library open chrome driver without the need to geckodriver and get up to date driver for you.

Related