I have some data read in from a poorly formatted pdf table, where the cells sometimes spans across several pages. This has left me with a dataframe that looks similar to this:
company_name <- c("company_a", NA, "company_a", "company_b", "company_b", NA)
text <- c("some_text", "text that should be in the above cell","some_text", "some_text", "some_text","text that should be in the above cell")
more_text <- c("some_text", "text that should be in the above cell", "some_text", "some_text", "some_text","text that should be in the above cell")
df <- data.frame(company_name, text, more_text)
| company_name | text | more_text |
|---|---|---|
| company_a | some_text | some_text |
| NA | text that should be in the above cell | text that should be in the above cell |
| company_a | some_text | some_text |
| company_b | some_text | some_text |
| company_b | some_text | some_text |
| NA | text that should be in the above cell | text that should be in the above cell |
How could I merge the rows that have a missing value where "company_name" should be, so it looks more like this and also loop it over all rows that start with NA:
| company_name | text | more_text |
|---|---|---|
| company_a | some_text + text that should be in the above cell | some_text + text that should be in the above cell |
| company_a | some_text | some_text |
| company_b | some_text | some_text |
| company_b | some_text + text that should be in the above cell | some_text + text that should be in the above cell |
I've tried the unheadr package, but I can't seem to figure out the correct function to use.
Edit: re-did the example for more clarity