I am doing web(twitter) crawling. (The reason we do not use api is to get past 7 days old data.)
This is a site where you can scroll and view your data.
If this takes a bigger range, "ConnectionResetError: [WinError 10054] The current connection has been forcibly disconnected by the remote host" error occurs.
Why am I not running?
Do you have a range limit on the Twitter page?
When I look at the other codes, I wonder if it works even if the range is set to 10000.
#python3.6, windows
import requests
import time
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup
browser = webdriver.PhantomJS('C:\phantomjs-2.1.1-windows/bin/phantomjs')
url = u'https://twitter.com/search?f=tweets&vertical=default&q=%EC%A0%95%EB%B6%80%20since%3A2017-07-20%20until%3A2017-07-21&l=ko&src=typd'
browser.get(url)
time.sleep(1)
body = browser.find_element_by_tag_name('body')
browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")
for _ in range(10000):
browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")
time.sleep(0.1)
tweets=browser.find_elements_by_class_name('tweet-text')
wfile = open("test.txt", mode='w', encoding='utf8')
data={}
i = 1
for i, tweet in enumerate(tweets):
data['text'] = tweet.text
print(i, ":", data)
wfile.write(str(data) +'\n')
i += 1
wfile.close()