I have a large number of text files containing rows of data and extra information. I would like to loop through the files and combine the data of interest into a single dataframe.
Each text file contains random information (rows of sentences, ect..) that i dont care about before and after actual data, but the exact number of rows before and after the data are highly inconsistent across text files. Thus, I cannot use typical arguments like skip or n_max to specify the rows I wish to read.
The only consistent patterns in the files are:
- before the data starts, there is a row containing the column headers for the data, and a row containing series of dashes
- When the data ends, there is a blank row, followed by a row that starts with the word "finished", and another row of dashes
examples of the data files are below: File 1:
i dont care
not important
this row is not important
Header starts on the next row
Index Date Time DP1 Name
--------------------------------------------------
1 07-20-22 17:48:06 3792123 machine 3
2 07-20-22 17:38:06 379211 machine 3
3 07-20-22 19:28:06 machine
4 07-20-22 19:48:06 379245 machine
5 07-20-22 17:58:06 37921 machine 2
--------------------------------------------------
finished blah blah
more rows
File2:
i dont care about this row and would like to remove it
Header starts on the next row
Index Date Time DP1 Name
--------------------------------------------------
1 07-20-22 17:48:06 machine 4
2 07-20-22 17:38:06 machine 8
3 07-20-22 19:28:06 machine
10 07-20-22 19:48:06 379245 machine
11 07-20-22 17:58:06 37921 machine 10
--------------------------------------------------
finished blah blah
Note the following:
- possible blanks in the fourth column
DP1 - inconsistent spacing between datapoints
- unpredictable lengths of words and sentences above and below the "data"
- the
Namecolumn could be one word or contain a space between a word and a number
Is there a way to use consistent patterns to loop through these files and compile the data of interest without having to touch the raw text files? My interest in this is not only for speed in manipulating the data, but to remove human-induced error and lack of transparency that could occur if I manipulate the raw files by hand.