I am trying to develop an application that needs to do web scraping. It is a complex scraping so I need a powerful tool such as Selenium.
The problem I have is that after few calls I start having errors:
Process running mem=544M(100.3%)
What I tried at the moment:
- setting start and max memory of JVM to 64m! (yes I don't need to much memory for my app, but did nothing). This option remained active also with the following attempts.
- creating and quitting WebDriver (Chrome) at every usage: it seems to work, but it is basically useless because it needs 2 seconds just to retrieve one instance (poor solution)
- pooling some instances in async just to avoid those 2 seconds: after 3 instances application is unable to create new instance with errors such as missing file blah blah blah... and anyway because of the scraping I need to do a lot of things everytime that I could avoid if I could use the same WebDriver more than once
- creating a wrapper of WebDriver that
quitand recreate the wrapped one after a number of usages: after 20 mins the errors appeared again
I am wondering if there is a solution that is alternative to 2.
I read a lot of answers saying that you need to quit/close, but I have just one window open and reusing it would be more effective than creating a new instance everytime I have to do something...
Do you have any suggestion?
Tech data:
- Java 11
- Selenium 3.141.59
- Heroku free tier