I am doing a task that requires scraping. I have a dataset with ids and for each id i need to scrape some new information. This dataset has around 4 million rows. Here is my code:
import pandas as pd
import numpy as np
import semanticscholar as sch
import time
# dataset with ids
df = pd.read_csv('paperIds-1975-2005-2015-2.tsv', sep='\t', names=["id"])
# columns that will be produced
cols = ['id', 'abstract', 'arxivId', 'authors',
'citationVelocity', 'citations',
'corpusId', 'doi', 'fieldsOfStudy',
'influentialCitationCount', 'is_open_access',
'is_publisher_licensed', 'paperId',
'references', 'title', 'topics',
'url', 'venue', 'year']
# a new dataframe that we will append the scraped results
new_df = pd.DataFrame(columns=cols)
# a counter so we know when every 100000 papers are scraped
c = 0
i = 0
while i < df.shape[0]:
try:
paper = sch.paper(df.id[i], timeout=10) # scrape the paper
new_df = new_df.append([df.id[i]]+paper, ignore_index=True) # append to the new dataframe
new_df.to_csv('abstracts_impact.csv', index=False) # save it
if i % 100000 == 0: # to check how much we did
print(c)
c += 1
i += 1
except:
time.sleep(60)
The problem is that the dataset is pretty big and this approach is not working. I left it working for 2 days and it scraped around 100000 ids, and then suddenly just froze and all the data that was saved was just empty rows. I was thinking that the best solution would be to parallelize and batch processing. I never have done this before and I am not familiar with these concepts. Any help would be appreciated. Thank you!
