Scrapy CoinMarketCap: How can I scrape and get information from the first page, scroll through others and aggregate information based on a filter?

Viewed 52

I'm a newbie in scrapy and python, although they have an API I'm doing a project to study the coinmarketcap site. I have some problems.

Question 1 - How to save the information of the first page and the pages that I'm going to go through in the same file (scrapy crawl -O cmc.csv)?

Result I want a csv like this: | Name | Price | Watchlists | | ------ | ------ |------------| | First | row | 3.054.863 | | Second | row | 2.056.312 |

I want to go through the first page: https://coinmarketcap.com/ Inform how many coins of the rank I want to analyze (or all). And get some information.

import scrapy


class CmcSpider(scrapy.Spider):
    name = 'cmc'
    start_urls = ['http://coinmarketcap.com/']

    def parse(self, response):
        for coin in response.css('td:nth-child(3)'):
            name_coin = coin.css('td:nth-child(3) p::text').get()
            price = coin.css('td:nth-child(4) span::text').get()
            vol_24h = coin.css('td:nth-child(5) span::text').get()

            yield {
                "name": name_coin, "price": price, "vol_24H": vol_24h
            }

        for item in response.css("tbody tr"):
            url = item.css("td:nth-child(3) a::attr(href)").get()
            yield scrapy.Request(url=f'http://coinmarketcap.com/{url}', callback=self.parse_currency)

    def parse_currency(self, response):
        name_second = response.css('.sc-103s2w8-0.eAmmwa span::text').get()

        yield {
            'name': name_second,
        }

Code under construction:

Question 2: I'm not able to separate the information "on watchlists". My snippet: wathclists = response.css('.bILTHz').get() Returns:

<div display="flex" style="flex-wrap:wrap" class="sc-16r8icm-0 bILTHz"><div class="namePill namePillPrimary">Rank #1</div><div class="namePill " style="text-transform:capitalize">Coin</div><div class="namePill">On 3,303,992 watchlists</div></div>

I'm not able to access only the watchlist information.

Question 3 - when i run the code on scrapy shell it only fetches the first 16 items

0 Answers
Related