The page I'm trying to scrape loads more items as you scroll down. I know how to make playwright scroll, I'm using a coroutine like this currently
PageCoroutine("evaluate", "window.scrollBy(0, document.body.scrollHeight)")
which when paired with another
wait_for_selector
coroutine where the selector is the id of the last item on the page works. My issue is the last item changes often and so I can't rely on it.
How can I tell scrapy/playwright to just keep scrolling until the bottom without needing to identify an element at the bottom?
Thanks
Here is my spider incase it helps:
import scrapy
from scrapy_playwright.page import PageCoroutine
class MySpider(scrapy.Spider):
name = 'my_spider'
def start_requests(self):
yield scrapy.Request(
'my-url',
meta={
'playwright': True,
'playwright_include_page': True,
'playwright_page_coroutines': [
PageCoroutine("evaluate", "window.scrollBy(0, document.body.scrollHeight)"),
PageCoroutine("wait_for_selector", "{item_at_bottom}"),
]
}
)
async def parse(self, response):
pass
# parses my content