My code that I usually use for getting simple HTML data into a dataframe is returning the IndexError: list index out of range message when I am trying to read in listed company data from Nasdaq and cant see where the issue is.
I would really appreciate some help.
from bs4 import BeautifulSoup
import pandas as pd
import requests
url = 'https://www.nasdaq.com/market-activity/stocks/screener'
headers = {"user-agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/85.0.4183.121 Safari/537.36"}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')
tables = soup.find_all('table', rules = 'all')
table = str(tables[0]) #cast table to string
df = pd.read_html(table, skiprows=2, flavor='bs4')[0]
print(df.head())
This is the 1st HTML table I want to read into a dataframe...out of 406. Each page has 20 rows of company data.
