Python requests_html render runs forever on certain URLs

Viewed 1785

I am trying to write a simple script that given an arbitrary URL will return the title tag of that website. Because many of the URLs I want to resolve need to have JavaScript enabled, I need to use something like requests_html's render function to do this. However, I have encountered an issue with the library where the example URL below never terminates. I have tried the timeout arg of the render call and it did not work. Can anyone help me figure out how to get this to timeout properly or some other work around to make sure it doesn't get stuck?

This is my current code that does not terminate (it gets stuck on the render call):

from requests_html import HTMLSession

session = HTMLSession()
r = session.get('http://shan-shui-inf.lingdong.works/')
# render with JS
r.html.render(sleep = 1, keep_page=True)
# Also does not work: r.html.render(sleep = 1, keep_page=True, timeout = 3)


title = r.html.find('title', first=True).full_text

I have already tried solutions like: Timeout on a function call and Python timeout decorator which still did not timeout strangely enough.

NOTE: I am using Python 3.7.4 64-bit on Windows 10.

2 Answers

I would suggest to put r.session.close() at last. This worked for me.

Ok I'm quite late here,
This is what I've done:

pip install -U pyppeteer

(pip installed the 0.2.6 version for me)

Then it worked somehow

(unrelated)
If you want the Chromium browser to appear on the screen you'll need to change requests_html.py (somewhere in site-packages)'s 714th line, headless=True -> headless=False

Related