Download a page from the command line using tools like selenium, curl, wget, etc

Viewed 314

I have been trying in the past several days to automate the saving from the command line of the following page as a table or a text file, using tools like wget or curl, but with no luck. The problems come also from the fact that the url is masked. I was wondering if there was the possibility to do it using tools like selenium.

https://www.hkex.com.hk/Market-Data/Securities-Prices/Equities?sc_lang=en

Before saving the page as a table, basically two operations need to be done:

a) Click once the '20 Items' on the lower right and bring it to '100 Items'

b) Click 10 times on 'LOAD MORE' link on the lower middle of the page, in order to increase the number of displayed items

I would greatly appreciate any suggestion on how to resolve this task. Thanks for your consideration.

1 Answers

The easiest way would be to use playwright - https://playwright.dev/python/docs/intro

Then record button clicks - playwright codegen

Save page contents/Export to pdf

You could save the script and then rerun it later - python hkex.py

example script

from playwright.sync_api import Playwright, sync_playwright
def run(playwright: Playwright) -> None:
    browser = playwright.chromium.launch(headless=False)
    context = browser.new_context()
    # Open new page
    page = context.new_page()
    # Go to https://www.hkex.com.hk/Market-Data/Securities-Prices/Equities?sc_lang=en
    page.goto("https://www.hkex.com.hk/Market-Data/Securities-Prices/Equities?sc_lang=en")
    # Click #accpet_cookie_btn
    page.click("#accpet_cookie_btn")
    # Click text=20 Items
    page.click("text=20 Items")
    # Click :nth-match(:text("Items"), 4)
    page.click(":nth-match(:text(\"Items\"), 4)")
    # Click text=LOAD MORE
    page.click("text=LOAD MORE")
    # Click text=LOAD MORE
    page.click("text=LOAD MORE")
    # Click text=LOAD MORE 100 Items 20 Items 50 Items 100 Items
    page.click("text=LOAD MORE 100 Items 20 Items 50 Items 100 Items")
    # Click text=LOAD MORE
    page.click("text=LOAD MORE")
    # Click text=LOAD MORE
    page.click("text=LOAD MORE")
    # Click text=Equity Overview LIST OF SECURITIES EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs
    page.click("text=Equity Overview LIST OF SECURITIES EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs")
    # Click .back_to_top_btn
    page.click(".back_to_top_btn")
    # assert page.url == "https://www.hkex.com.hk/Market-Data/Securities-Prices/Equities?sc_lang=en#hkex_page_header"
    # Click text=Equity Overview LIST OF SECURITIES EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs
    page.click("text=Equity Overview LIST OF SECURITIES EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs")
    # Click #lhkexw-equities div:has-text("FILTERS EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs DEBT SECURITIES Filters EQ")
    page.click("#lhkexw-equities div:has-text(\"FILTERS EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs DEBT SECURITIES Filters EQ\")")
    # Click #lhkexw-equities div:has-text("FILTERS EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs DEBT SECURITIES Filters EQ")
    page.click("#lhkexw-equities div:has-text(\"FILTERS EQUITIES ETPs DWs INLINE WARRANTS CBBCs REITs DEBT SECURITIES Filters EQ\")")
    # Click text=LOAD MORE
    page.click("text=LOAD MORE")
    # ---------------------
    context.close()
    browser.close()
with sync_playwright() as playwright:
    run(playwright)
Related