range limit on the web page (Twitter)?

Viewed 109

I am doing web(twitter) crawling. (The reason we do not use api is to get past 7 days old data.)

This is a site where you can scroll and view your data.

If this takes a bigger range, "ConnectionResetError: [WinError 10054] The current connection has been forcibly disconnected by the remote host" error occurs.

Why am I not running?

Do you have a range limit on the Twitter page?

When I look at the other codes, I wonder if it works even if the range is set to 10000.

#python3.6, windows
import requests
import time
from selenium import webdriver
from selenium.webdriver.common.keys import Keys
from bs4 import BeautifulSoup

browser = webdriver.PhantomJS('C:\phantomjs-2.1.1-windows/bin/phantomjs')
url = u'https://twitter.com/search?f=tweets&vertical=default&q=%EC%A0%95%EB%B6%80%20since%3A2017-07-20%20until%3A2017-07-21&l=ko&src=typd'


browser.get(url)
time.sleep(1)

body = browser.find_element_by_tag_name('body')
browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")

for _ in range(10000):
    browser.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    time.sleep(0.1)

tweets=browser.find_elements_by_class_name('tweet-text')

wfile = open("test.txt", mode='w', encoding='utf8')
data={} 
i = 1
for i, tweet in enumerate(tweets):
    data['text'] = tweet.text
    print(i, ":", data)
    wfile.write(str(data) +'\n')
    i += 1
wfile.close()
0 Answers
Related