I wrote a multithreaded web crawler under Windows. The libraries that I used were requests and threading. I found the program became slower and slower after running for some time (about 500 pages). When I stop the program and run again, the program speeds up again. It seems that there are many pending connections, causing the slowdown. How should I manage the problem?
My code:
import requests, threading,queue
req = requests.Session()
urlQueue = queue.Queue()
pageList = []
urlList = [url1,url2,....url500]
[urlQueue.put(i) for i in urlList]
def parse(urlQueue):
try:
url = urlQueue.get_nowait()
except:
break
try:
page = req.get(url)
pageList.append(page)
except:
continue
if __name__ == '__main__':
threadNum = 4
threadList = []
for i in threadNum:
t = threading.Thread(target=(parse),args=(urlQueue,))
threadList.append(t)
for thread in threadList:
thread.start()
for thread in threadList:
thread.join()
I searched for the problem. An answer told that it was the reuse and recycling problem of TCP under Linux. I don't understand that answer very well. The answer is below. I translated the answer from the Chinese.
- Type command in Linux shell:
netstat -n | awk '/^tcp/ {++S[$NF]} END {for(a in S) print a, S[a]}' - Found the
TIME_WAITis nearly 2W. So, there must be many TCP connections. - Use the following code to set the reuse time and recycling time, respectively of TCP:
echo "1" > /proc/sys/net/ipv4/tcp_tw_reuse,echo "1" > /proc/sys/net/ipv4/tcp_tw_recycle
That answer seems correct. It should be a network problem. How should I solve this under Windows.