I want to scrape data off of this website:
https://www.gurufocus.com/stock/AAPL/
This is the current version of the webscraper:
import pandas as pd
import requests
from bs4 import BeautifulSoup
import re
ls=['Ticker','PE Ratio','Gross Margin %','Debt-to-EBITDA','GF Value Rank','Financial Strength','Profitability Rank']
symbols = ['AAPL', 'TSLA']
df = pd.DataFrame(columns=ls)
for t in symbols:
req = requests.get("https://www.gurufocus.com/stock/"+t)
if req.status_code !=200:
continue
soup = BeautifulSoup(req.content, 'html.parser')
scores = [t]
for val in ls[1:]:
scores.append(soup.find('a', string=re.compile(val)).find_next('td').text)
df.loc[len(df)] = scores
df
And this is the output that I get:
The normal ratios are correctly obtained, but the GF Value Rank, the Financial Strenght and Profitability Rank couldn't be obtained with the program above.
I inspected the webcode of the gurufocus website and came across this div section with the id="financial-strength" and id="profitability", but I'm not sure how to extract the scores from this information.
As far as the GF Value Rank is concerned, I only found a span section that covers that score, but nothing like a javascript td entry or something similar.
How do I need to change my code to obtain the last three scores in my table?
