I am writing a parser for a custom file format constituted of an undetermined number of data structures spanning over multiple lines. The lines constituting each individual structure are garanteed to be contiguous, but can appear in any order. In order to be able to assemble the structure, the offset of the data in the structure on each line is written at the begginning of said line.
I wrote a lexer that produces 3 types of tokens : data_offset, data, end_of_line. But that means it is now up to my parser to discard data_offset and end_of_line tokens, and rearrange the lines, which I feel like is not appropriate.
Should I instead write a lexer that produce only structure tokens, representing the structures stripped of the offset and EOL and reassembled into the right order?
Edit:
If anyone happens to stumble upon this post, What I ended up doing is a classic lexer/parser which simply tokenize file and builds an AST, then I implemented an "analyser" that can work on the parsed data, rearrange it etc... "Every man to his trade"