How does Scrapy-Splash implement Proxy Profiles?

Viewed 1208

I'm having some problems with using Scrapy-Splash together with an HTTP proxy (see "500 Internal Server Error" when combining Scrapy over Splash with an HTTP proxy), even when I try to set a proxy profile following http://splash.readthedocs.io/en/latest/api.html#proxy-profiles.

In order to understand better what is going on, I was looking for the part of the Scrapy-Splash source code, https://github.com/scrapy-plugins/scrapy-splash, which parses the proxy host and port specified in the .ini file in /etc/splash/proxy-profiles.

However, searches for "proxy" or ".ini" in the repository didn't yield any results. Can someone explain to me how the proxy profiling is implemented in Scrapy-Splash?

1 Answers

First, the Scrapy-Splash proxy setting is in /etc/splash/proxy-profiles, but if you are running splash in a container, you can map the host proxy profile to the container by -v, eg:

sudo docker run -p 8050:8050 -v /etc/splash/proxy-profiles:/etc/splash/proxy-profiles scrapinghub/splash

Second, when visiting the url through splash, a proxy parameter is need if proxy profile name is not default.ini, eg:

localhost:8050/render.html?url=http://target.com?wait=1&timeout=2&proxy=filename
Related