I've been trying to use tabulizer to avoid hardcoding parsing that could possibly change with the next report. I was wondering if you all might have better ideas.
library(tabulizer)
library(tidyverse)
who <- "https://www.who.int/docs/default-source/coronaviruse/situation-reports/20200309-sitrep-49-covid-19.pdf"
page1 <- tabulizer::extract_tables(who, pages = 4, output = "data.frame") %>%
as.data.frame() %>%
slice(5:n()) %>%
select(-`X.1`)
page2 <- tabulizer::extract_tables(who, pages = 5, output = "data.frame") %>%
as.data.frame() %>%
rbind(colnames(.))
page3 <- tabulizer::extract_tables(who, pages = 6, output = "data.frame") %>%
as.data.frame() %>%
rbind(colnames(.))
colnames(page2) <- colnames(page1)
colnames(page3) <- colnames(page1)
dat <- page1 %>% rbind(page2) %>% rbind(page3)
If you run this you'll notice the region and totals need to be removed but the page split and tall rows are where I'm having trouble.