I'm a newbie in scrapy and python, although they have an API I'm doing a project to study the coinmarketcap site. I have some problems.
Question 1 - How to save the information of the first page and the pages that I'm going to go through in the same file (scrapy crawl -O cmc.csv)?
Result I want a csv like this: | Name | Price | Watchlists | | ------ | ------ |------------| | First | row | 3.054.863 | | Second | row | 2.056.312 |
I want to go through the first page: https://coinmarketcap.com/ Inform how many coins of the rank I want to analyze (or all). And get some information.
import scrapy
class CmcSpider(scrapy.Spider):
name = 'cmc'
start_urls = ['http://coinmarketcap.com/']
def parse(self, response):
for coin in response.css('td:nth-child(3)'):
name_coin = coin.css('td:nth-child(3) p::text').get()
price = coin.css('td:nth-child(4) span::text').get()
vol_24h = coin.css('td:nth-child(5) span::text').get()
yield {
"name": name_coin, "price": price, "vol_24H": vol_24h
}
for item in response.css("tbody tr"):
url = item.css("td:nth-child(3) a::attr(href)").get()
yield scrapy.Request(url=f'http://coinmarketcap.com/{url}', callback=self.parse_currency)
def parse_currency(self, response):
name_second = response.css('.sc-103s2w8-0.eAmmwa span::text').get()
yield {
'name': name_second,
}
Code under construction:
Question 2: I'm not able to separate the information "on watchlists". My snippet:
wathclists = response.css('.bILTHz').get()
Returns:
<div display="flex" style="flex-wrap:wrap" class="sc-16r8icm-0 bILTHz"><div class="namePill namePillPrimary">Rank #1</div><div class="namePill " style="text-transform:capitalize">Coin</div><div class="namePill">On 3,303,992 watchlists</div></div>
I'm not able to access only the watchlist information.
Question 3 - when i run the code on scrapy shell it only fetches the first 16 items