Find a pattern in a stream of bytes read in blocks

Viewed 10

I have a stream of gigabytes of data that I read in blocks of 1 MB.

I'd like to find if one of the patterns PATTERNS = [b"foo", b"bar", ...] is present or not in the data (case insensitive).

Here is what I'm doing. It works but it is sub-optimal:

oldblock = b''
while True:
    block = data.read(1024*1024)
    if block == b'':
        break
    testblock = (oldblock + block).lower()
    for PATTERN in PATTERNS:
        if PATTERN in testblock:
            for l in testblock.split(b'\n'):  # display only the line where the 
                if PATTERN in l:              # pattern is found, not the whole 1MB block!
                    print(l)                  # note: this line can be incomplete if 
    oldblock = block                          # it continues in the next block (**)

Why do we need to search in oldblock + block? This is because the pattern foo could be precisely split in two consecutive 1 MB blocks:

[.......fo] [o........]
block n     block n+1

Drawback: it's slow to concatenate oldblock + block and to perform the search nearly in double.

We could use testblock = oldblock[-max_len_of_patterns:] + block, but there is surely a more canonical way to address this problem, as well as the side-remark (**).

How to do a more efficient pattern search in data read by blocks?

1 Answers
  1. if patern matched, try to use "break;" word inside "for" body for breaking execution of already useless code
  2. and use {...} for start and finish "for" loop body like:

for (...) { if match(PATTERN) break; }

Related