Parsing Farsi strings using pdf_text in R

Viewed 127

I am attempting to parse this pdf file from the Afghanistan Independent Election Commission (listing polling centers planned for elections held on Oct 20) into a csv file using R. I was able to successfully do this for the English-language version of the same list, but my objective here is to extract the Farsi (Dari)-language names of each polling center, so as to be able to join those centers with other voter registration data in which the polling center code has been omitted but I have a center name. (I am not 100% certain that names will be consistently rendered across datasets but it's my best hope for a match at this point.)

In order for that join I would need to be able to extract the exact polling center Farsi name strings from the polling center list. Unfortunately, I am running into issues using the pdftools library for this. I am not clear if this is an issue with the pdf file's encoding, or a broader issue with Farsi-language reading in pdftools (and I don't read the language myself, though I can read the characters). The Farsi strings are in some cases being broken up by extra spaces that are not visually evident in the pdf, which is confounding my attempts to correctly parse them into columns and would presumably prevent an accurate match even if I were able to force them into the correct columns. (I am having the same issue reading the voter registration files, so either the underlying issue is consistent across files or my mistaken approach is.)

Sample code:

library(pdftools)
# import

pc_import <- pdf_text("http://www.iec.org.af/pdf/pclist-2018/allpc1397.pdf")
pdf_text <- toString(pc_import)
pdf_text <- read_lines(pdf_text)

# strip headers and footers
header_row_1 <- grep(trimws(pdf_text[1]), pdf_text)
footer_row <- grep("Page ", pdf_text)
pdf_text <- pdf_text[- c(header_row_1, header_row_1 + 1, header_row_1 + 2, footer_row)]

# convert to a dataframe
data <- Reduce(rbind, strsplit(trimws(pdf_text), "\\s{2,}"))
rownames(data) <- 1:dim(data)[1]
colnames(data) <- c("pc_location_dari", "pc_name_dari", "pc_code", "district_code", "district_name_dari", "province_code", "province_name_dari", "id")
data <- as.data.frame(data)

The erroneous string splitting issue can be observed in pdf_text and data rows 4-6.

I've also attempted to process the polling center file with Tabula, but while the output there to be slightly more accurate, it is still producing similar split string issues in some cases, which makes me think there may be some issue in the underlying file encoding. And suggestions on how to approach this would be appreciated.

0 Answers
Related