I was scraping this site with the following code:
import requests
from bs4 import BeautifulSoup
url = "https://www.pro-football-reference.com/teams/buf/2021_injuries.htm"
r = requests.get(url)
stats_page = BeautifulSoup(r.content, features="lxml")
table = stats_page.findAll('table')[0] #get FIRST table on page
for player in table.findAll("tr"):
print([i.getText() for i in player.findAll("td")])
The output is:
[]
['', 'IR', 'IR', 'IR', 'IR', 'IR', 'IR', 'IR']
['', 'Q', '', '', '', '', '', '']
['', '', '', '', 'Q', '', '', '']
['', '', '', 'O', '', '', '', 'IR']
['', '', 'Q', '', '', '', '', '']
['', '', '', 'Q', '', '', '', '']
['', '', '', '', 'Q', '', '', '']
['O', 'Q', '', '', '', '', '', '']
['', '', '', '', 'Q', '', '', '']
['', 'Q', '', 'Q', '', '', '', '']
['', '', '', 'O', '', '', '', '']
['Q', '', '', '', '', '', '', '']
['', 'IR', 'IR', 'IR', 'IR', 'IR', 'IR', 'IR']
['', '', 'Q', '', '', '', '', '']
['', 'IR', 'IR', 'IR', 'IR', '', '', '']
This is clearly the output I would expect from the 2nd table on the page, "Team Injuries", rather than the 1st table on the page, "Week 10 injury report". Any idea why BeautifulSoup is seemingly ignoring the first table on the page?