I have a function which does the following:
- open chrome
- scrape data
- close chrome
Within this function I create a dataframe that should include all the elements in a list: however, an error occurs (captcha) after each element (after the first one), so I need to create a document (csv file) and update it every iteration.
What I need is doing this:
- check the first element x1 in a list
- open chrome
- scrape data for the first element x1 in the list
- close chrome (because an error occurs due to the captcha)
- create a csv file with fields for the first element x1 in the list
- remove the first element x1 from that list
- check the second element x2 in a list if it is not already in the csv file (Source column)
- open chrome
- scrape data for the second element x2 in the list
- close chrome (because an error occurs due to the captcha)
- update the csv file with fields for the second element x2 in the list
- etc.
Let's say the list has 5 elements. My function looks like:
def my_func(df):
chrome_options = webdriver.ChromeOptions()
renew=[]
tags=[]
query=df['Source'].unique().tolist()
options = {'my:options'}
driver=webdriver.Chrome('my_path',chrome_options=chrome_options)
driver.maximize_window()
frame_dict={}
for x in query:
print(x)
response=driver.get('website/'+x)
try:
wait = WebDriverWait(driver, 30)
time.sleep(randrange(5))
driver.execute_script("window.scrollTo(0, 1000)")
try:
# some code
renew.append(w_renew)
except:
renew.append("Data not available")
# Tags
try:
# some code
tags.append(tag)
except:
tags.append("Data not available")
except:
print("\n!!! Error !!!")
break
# Create dataframe
print("\n")
frame_dict.update({'My_List': x, 'Col1': col_1, 'Col2': col_2})
print(frame_dict)
df=pd.DataFrame.from_dict(frame_dict)
driver.quit()
df.to_csv("path/my_file.csv")
return df
Currently my code above is not removing any item from the list and it runs only once, before the error occurs (except print Error). The code opens Chrome, then it closes Chrome, but data are overwritten after I re-run the code (starting from the same element just scraped, since it is not removed from the list; I would need to start from the next element). I think a loop cycle could fix the problem of iterations, but I do not know how to create/update the df/csv file and moving to other elements in the list, finally stopping the process when all the elements in the list are correctly scraped and included in the csv file. I hope you can help me.