Capture a captcha image

Viewed 301

I am trying to scrape data from the Peruvian Electronic System for Government Procurement and Contracting (SEACE) (using RSelenium) and I have succeeded until I try to capture the URL from the Captcha image. The problem that I encounter is that the link for the captcha has the extension "dynamiccontent.properties.xhtml" (See next screenshot), but not a "JPEG", "JPG" or "PNG" extension.

Screenshot from SEACE

I would like to get the URL from the captcha image in one of these extensions (JPEG, JPG or PNG) using R, any suggestions? Thanks!

1 Answers

You CAN get the captcha image using Rselenium. But you require some image processing along with it. Since the captcha is generated dynamically, you need to take a screenshot of the page, then crop the image such that only the captcha is left. (try playing with the dimension argument of the cropping function to get it just right)

While doing so, make the window size large so that the resolution is good.

For cropping the image, you will need some trial and error. You can use packages imager or magick for the cropping bit.

library(RSelenium)
library(magick)
library(dplyr)
url<-"https://prodapp2.seace.gob.pe/seacebus-uiwd-pub/buscadorPublico/buscadorPublico.xhtml"

#### Selenium server
## For Firefox

rd<- rsDriver(browser = "firefox",port = 4581L)
remDr <- rd$client
remDr$open()


##Open url
remDr$navigate(url)

remDr$setWindowSize(2000,1600)

remDr$screenshot(display = FALSE, file = "captcha.png")


final_cap<-image_read("captcha.png") %>% 
  image_crop(.,"200X100+500+500") 

plot(final_cap)

Please do note that I don't support any illegal captcha breaking - they were designed to keep robots out!

Captcha Captured!

Related