How can I parse a txt file using multiple wildcards?

Viewed 104

The goal is to get just the updated numeric values along with their identifiers while ignoring everything else. I have a CDM.txt file with 1 line of information that gets updated every couple hours and the numeric values all change.

'MISS_DISTANCE': '18', 'MESSAGE': 'NEED HELP', 'MISS_DISTANCE': '105398', ETC, ETC

I would like to parse the text file and set the output to variable "Miss Distance". Using that example, I would need to get the following and ignore the rest.

'MISS_DISTANCE': '18' | 'MISS_DISTANCE': '105398'

Here is what I have so far.

with open('CDM.txt') as f:
lines = f.readlines()
reg = re.compile("'MISS_DISTANCE': '.+',")

Not sure what to do from there or if I am even doing the compile correctly. I added the .+ for multiple wildcards since the number can range between 1 to 99999 and a "," at the end because I want it to end there.

3 Answers

This opens two files, your file (which I call blankpaper.txt) and an out.txt file. The latter we will write to. If the file is not too large, we can read in the text of the whole file into memory and make text refer to it then use findall() from the re module to find all matches of the regular expression we compiled. findall returns a list of all such matches. We can then use this information however we like, for example, writing to a new file. The regular expression matches 'MISS_DISTANCE:' followed by 1 or more white-space characters (\s+) followed by a ' followed by one or more decimal digits (\d+)followed by a '.

import re

p = re.compile(r"'MISS_DISTANCE':\s+'\d+'")
with open("blankpaper.txt") as f, open("out.txt", "w") as f_out:
    text = f.read()
    m = re.findall(p, text)

    f_out.write(f"{m[0]} ")
    for i in range(1, len(m)):
        f_out.write(f"| {m[i]} ")

print(m)

Output of print(m)

["'MISS_DISTANCE': '18'", "'MISS_DISTANCE': '105398'"]

out.txt

'MISS_DISTANCE': '18' | 'MISS_DISTANCE': '105398' 

Since the number is in the range 1 to 99_999, you might want to specify the number of digits :

reg = re.compile("'MISS_DISTANCE': '\d{1,5}',")

Don't forget that the last item is not followed by a comma ! You should remove the comma in the regular expression, and add it after your match.

def _rows():
    with open("your.txt", 'r') as _file:
        text = _file.read()
        for item in text.split(', '):
            _, sep, word = item.partition(': ')
            if sep and word.strip('\'').isdecimal():
                yield item


print(*list(_rows()), sep=', ')

Maybe this is what you need. Using str.partition and str.isdecimal instead of re.search for better performance.

Related