I have several hundred documents of multiple pages with text which I extract from the web. I am trying to extract all country names which appear in the text plus e.g. 10 words before and after each match/country name.
So far i made use of the countrycode package which contains a dataframe codelist which contains names of basically all countries (as well as reprex versions of these names).
I pulled these names into one long vector, collapsed it with the OR sign "|" as a separator and use the resulting string as regex in stringr's str_extract_all. The approach by and large works, however, it is quite slow for my liking.
Checking only the first 1500 characters of one single document takes 40 seconds on my (average) laptop (16 gig, i5-8625U). Running the script on all my documents will probably take more than a day.
Hence, my question - how can I increase the speed of extracting country names and their neighboring words? Since this is a task I am regularly requiring I am happy also to take up other tools (Julia? Python?).
library(tidyverse)
library(countrycode)
library(rvest)
library(tictoc)
df_main <-c("http://docstore.ohchr.org/SelfServices/FilesHandler.ashx?enc=6QkG1d%2fPPRiCAqhKb7yhssh2tXWBbyLwahMw00Sn91Vx2CtIqUsxb8GHTWZijKzuMBGT4crbn9EZ8sIYVnIZ4cuyvgS8esYaZB%2bpIm6AjmiXFUTphoU1fK%2blmqzGzjoi8BFXLvYtOrKwuZsRFFTmN0umUat3QkQ%2bVArWqePYKE8%3d") %>%
enframe(name=NULL,
value="doc_link")
fn_get_doc_text <- function(doc_link) {
xml2::read_html(doc_link) %>%
rvest::html_text()
}
df_main <- df_main %>%
mutate(doc_text=map_chr(doc_link, possibly(fn_get_doc_text, otherwise="missing"))) %>%
mutate(doc_intro=str_sub(doc_text, end=1500) %>% str_squish())
df_country_patterns <- countrycode::codelist %>%
select(country.name.en) %>%
mutate(country_pattern=paste0("(\\S+\\s+){0,10}", country.name.en, "(\\s+\\S+){0,10}"))
vec_country_patterns <- df_country_patterns %>%
pull(country_pattern) %>%
paste(., collapse="|")
# stringr -----------------------------------------------------------------
tic()
df_main <- df_main %>%
mutate(countries_detected=str_extract_all(doc_intro, regex(vec_country_patterns,
ignore_case = F,
dotall = T)))
toc()
#> 41.34 sec elapsed
Created on 2020-09-20 by the reprex package (v0.3.0)