How to parse 380K html pages (best way in terms of performance)?

Viewed 91

What do I have:

  • 380,000 html pages (acticles from news portal from 2010 to 2017).
  • 65.19Gb (68,360,000Kb) total.
  • 170Kb per page (in average).
  • Server 1: 2.7GHz i5, 16Gb Ram, SSD, macOS 10.12, R version 3.4.0 (my laptop).
  • Server 2: 3.5GHz Xeon E-1240 v5, 32Gb, SSD, Windows Server 2012, R version 3.4.0.

What do I need:

  • I need to parse those html pages and get the data that looks like: ArticleDate | ArticleSection | ArcticleTitle | ArticleAuthor | ArticleContent| ... |

What do I do:

require(rvest)
files <- list.files(file.path(getwd(), "data"), full.names = TRUE, recursive = TRUE, pattern = "index")
downloadedLinks <- c()
for (i in 1:length(files)) {
  currentFile <- files[i]
  pg <- read_html(currentFile, encoding = "UTF-8")
  fileLink <- html_nodes(pg, xpath=".//link[@rel='canonical']") %>% html_attr("href")   
  downloadedLinks <- c(downloadedLinks, fileLink)  
}

I run this code on 40,000 pages and got this result:

  • Server 1: 800secs
  • Server 2: 1000secs

It means it will take 7600secs or 126mins to process 380,000 pages on Server 1 and 9500secs or 158mins on Server 2

Because of that I have few questions and hope that community will help me. I will be glad to hear any ideas, suggestions or critics.

  1. How can I improve my code above to reduce processing time?
  2. Why Server 2 (I mean real server) shows low performance comparing to Server 1 (my laptop)?
  3. Is there a way to order Server 2 to use more CPU and RAM (it is about 10% CPU usage and 300Mb RAM for Rgui.exe process)
0 Answers
Related