Simplified example of a text i have after importing with readlines:
text <- c("just", "stuff", "nothing", "interesting", "date", "06.05.2022",
"number", "1/3892", "adress", "north street 45", "name", "peter miller",
"just", "stuff", "nothing", "interesting", "date", "06.05.2022",
"number", "5/7283", "adress", "south street 11, fareaway", "west street 4",
"name", "john snow", "just", "stuff", "nothing", "interesting",
"date", "06.05.2022", "number", "7/112563", "adress", "island street 348",
"planet street 11, tortuga", "calvary road 9", "name", "hogson, michael",
"jobs, steve", "just", "stuff", "nothing", "interesting", "date",
"06.05.2022", "number", "2/1575", "adress", "bowland road 2, mexiko",
"name", "michael myers", "terry jones", "olivia wilde", "just",
"stuff", "nothing", "interesting", "date", "06.05.2022", "number",
"1/93375", "adress", "sunset boulevard", "name", "harrison ford")
The same pattern is always repeating, I would like to have a dataframe like this:
| date | number | adress | name |
|---|---|---|---|
| 06.05.2022 | 5/7283 | south street 11, fareaway, west street 4 | john snow |
| 06.05.2022 | 7/112563 | island street 348, planet street 11, tortuga, calvary road 9 | hogson, michael, jobs, steve |
There is always exact one date, one number, but one or more adresses and one or more names. "just stuff nothing interesting" is also always the same and can reliably be used the detect the end of the names.
I guess this could be achieved with loops, but I gave up on trying. Or is there a function which handles such irregularities? (not even sure if length is the right word for it, I hope it is clear what I mean...)