Regex for splitting document

Viewed 123

I would like to split individual PDF documents from a compound set from a file. So far it works for files which are structured: %PDF- .... %%EOF ... %PDF- .... %%EOF by using the following code:

REGEX_PDF = b'%PDF\-.+?%%EOF'
pdfDocuments = re.findall( REGEX_PDF, fileContent, re.DOTALL )

Now I need to change the software to also work with PDFs with extensions. This results in a file structure like this: %PDF- .... %%EOF ... %%EOF ... %%EOF ... %PDF- .... %%EOF. So I need to match to substrings from a PDF tag until the last %%EOF tag before the next PDF tag. My best guess is this:

REGEX_PDF = b'%PDF\-.+(?!%PDF\-).+%%EOF'

But it does not seem to work. Instead only 1 substring is matched from the 1st %PDF tag top the very last %%EOF tag. Does someone has an idea where the error is?

Thanks in advance, Thomas

1 Answers

You can rely on the `start delimiter here, and use

re.split(rb'(?!\A)(?=%PDF-)', fileContent)
re.findall(rb'%PDF-.*?(?=%PDF-|\Z)', fileContent, re.S)
re.findall(rb'%PDF-[^%]*(?:%(?!PDF-)[^%]*)*', fileContent)

See the regex #1 demo, regex #2 demo and regex #3 demo.

The (?!\A)(?=%PDF-) regex matches a location at a non-start position that is immediately followed with %PDF-.

The %PDF-.*?(?=%PDF-|\Z) pattern matches %PDF-, then any zero or more chars as few as possible up to the leftmost occurrence of %PDF- or end of string. %PDF-[^%]*(?:%(?!PDF-)[^%]*)* is almost the same, but it does not check if there is %PDF- on the right side (here, the (?=%PDF-|\Z) lookahead check is built ("married") into the .*? pattern).

Related