Get HTML table without id in Python Selenium

Viewed 52

I am scrapping the following page: https://www.sbs.gob.pe/app/pp/EstadisticasSAEEPortal/Paginas/TIPasivaDepositoEmpresa.aspx?tip=C

The first problem is that the table that I want has no id, so I use class name. image html

I wanna extract all the info from the selected table. The problem is that when I scrap it using selenium I do find the table but I can't access its body or childs.

Here is my python code:

driver = webdriver.Chrome( ChromeDriverManager().install() )
url = "https://www.sbs.gob.pe/app/pp/EstadisticasSAEEPortal/Paginas/TIPasivaDepositoEmpresa.aspx?tip=C"

wd = webdriver.Chrome( ChromeDriverManager().install() )
wd.maximize_window()
wd.get(url)

table_path = wd.find_elements_by_class_name("APLI_tabla")

innerhtml = table_path.get_attribute('innerHTML')     
 

The code above outputs the following error: AttributeError: 'list' object has no attribute 'get_attribute'

This is the specific table I want: table

1 Answers

It's good practice to wait for the element to be loaded (clickable) in page, before locating them. Something like below will grab all tables from the page and print them as dataframes, just pick the one you need. Setup is for linux, but you can adapt it to your own, just pay attention to imports, and to the code after the browser is getting the url:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

chrome_options = Options()
chrome_options.add_argument("--no-sandbox")
chrome_options.add_argument("window-size=1280,720")

webdriver_service = Service("chromedriver/chromedriver") ## path to where you saved chromedriver binary
browser = webdriver.Chrome(service=webdriver_service, options=chrome_options)

url = 'https://www.sbs.gob.pe/app/pp/EstadisticasSAEEPortal/Paginas/TIPasivaDepositoEmpresa.aspx?tip=C' 
browser.get(url)

tables = WebDriverWait(browser, 20).until(EC.presence_of_all_elements_located((By.TAG_NAME, "table")))
print(len(tables))
for table in tables:
    
#     print(table.get_attribute('outerHTML'))
    df = pd.read_html(table.get_attribute('outerHTML'))
    print(df)
    print('__________________')

This will print out all the tables as dataframes:

15
[]
__________________
[                                                           0
0  TASA DE INTERÉS PROMEDIO DEL SISTEMA DE CAJAS MUNICIPALES]
__________________
[                            0       1      2   3   4
0  Ingrese fecha: 2022  Junio     NaN    NaN NaN NaN
1  Ingrese fecha: 2022  Junio     NaN    NaN NaN NaN
2              Ingrese fecha:  2022.0  Junio NaN NaN,                             0       1      2   3   4
0  Ingrese fecha: 2022  Junio     NaN    NaN NaN NaN
1              Ingrese fecha:  2022.0  Junio NaN NaN,                 0     1      2   3   4
0  Ingrese fecha:  2022  Junio NaN NaN]
__________________
[                            0       1      2   3   4
0  Ingrese fecha: 2022  Junio     NaN    NaN NaN NaN
1              Ingrese fecha:  2022.0  Junio NaN NaN,                 0     1      2   3   4
0  Ingrese fecha:  2022  Junio NaN NaN]
__________________
Related