Well, I have been trying this for a few days. I have to say that I am very new to Python and I don't fully understand Javascript (for now), so maybe is something stupid but I can't figure it out.
I wan't to scrape the different tables on BSCScan, because of simplicity and because all of them are almost the same I will show the code of holders. I want to store it in a dict, then append to list of dicts and convert to data frame with pandas, with Address, quantity and percentage: This is the approach I made with html_requests:
from requests_html import HTMLSession
contract = "0x84c0160d55a05a28a034e1e6776f84c5995aba3a"
url = ("https://bscscan.com/token/" + contract + "#balances")
session = HTMLSession()
r = session.get(url)
r.html.render(sleep=3)
print(r.content)
holders = r.html.xpath('//td', first = True)
print(holders) #This returns None
This is the code I was making with requests_html, and with Selenium this is it:
driver_path = "/Users/XXX/Downloads/chromedriver"
driver = webdriver.Chrome(driver_path)
def get_top_holders(driver, holders):
list_holders = []
driver.get(holders)
tabla = driver.find_element_by_xpath('.//tbody').get_attribute('innerHTML') #I tried innerHTML just to see if it works in that way.
for td in tabla.find_elements_by_xpath('.//tr'):
name = td.find_element_by_xpath('.//td[2]/span/a').get_attribute('textContent')
quantity = td.find_element_by_xpath('.//td[3]').get_attribute('textContent')
percentage = td.find_element_by_xpath('//td[4]/text')
dict = {
'Addres': name,
'Cantidad': quantity,
'Porcentaje': percentage
}
os.system('clear')
print(dict)
driver.close()
list_holders.append(dict)
print(list_holders)
holders_tabla = pd.DataFrame(list_holders)
return holders_tabla
I have tried with Selenium, letting it render and trying to extract, but I can't iterate from tbody. I have tried with Beautiful Soup but I don't get it completely and someone recommend me requests_html but it is returning none.
First time asking, thanks in advance!