I'd like to extract certain numeric info from a bunch of PDFs. A sample is shown below, where the numeric info is positioned under the corresponding headings.
The strings corresponding to the above image (read in by pdftools::pdf_text()) is:
mystr <- ' Natural Dry\n Metric Tons @ Moisture or Metric Tons\n B.L. WEIGHT: 78,944 1.70% 77,601.952\n'
There are a lot of spaces and line breaks. Is it possible to extract the information under those headings?
My desired end result would be something like:
myresult <- tibble(
`Natural Metric Tons` = 78944,
Moisture = 1.7,
`Dry Metric Tons` = 77601.952
)
