Python remove Square brackets and extraneous information between them

Viewed 1291

I'm trying to handle a file, and I need to remove extraneous information in the file; notably, I'm trying to remove brackets [] including text inside and between bracket [] [] blocks, Saying that everything between these blocks including them itself but print everything outside it.

Below is my text File with data sample:

$ cat smb
Hi this is my config file.
Please dont delete it

[homes]
  browseable                     = No
  comment                        = Your Home
  create mode                    = 0640
  csc policy                     = disable
  directory mask                 = 0750
  public                         = No
  writeable                      = Yes

[proj]
  browseable                     = Yes
  comment                        = Project directories
  csc policy                     = disable
  path                           = /proj
  public                         = No
  writeable                      = Yes

[]

This last second line.
End of the line.

Desired Output:

Hi this is my config file.
Please dont delete it
This last second line.
End of the line.

What i have tried based on my understanding and re-search:

$ cat test.py
with open("smb", "r") as file:
  for line in file:
    start = line.find( '[' )
    end = line.find( ']' )
    if start != -1 and end != -1:
      result = line[start+1:end]
      print(result)

Output:

$ ./test.py
   homes
   proj
12 Answers

with one regex

import re

with open("smb", "r") as f: 
    txt = f.read()
    txt = re.sub(r'(\n\[)(.*?)(\[]\n)', '', txt, flags=re.DOTALL)

print(txt)

regex explanation:

(\n\[) find a sequence where there is a linebreak followed by a [

(\[]\n) find a sequence where there are [] followed by a linebreak

(.*?) remove everything in the middle of (\n\[) and (\[]\n)

re.DOTALL is used to prevent unnecessary backtracking


!!! PANDAS UPDATE !!!

The same solution with the same logic can be carried out with pandas

import re
import pandas as pd

# read each line in the file (one raw -> one line)
txt = pd.read_csv('smb',  sep = '\n', header=None)
# join all the line in the file separating them with '\n'
txt = '\n'.join(txt[0].to_list())
# apply the regex to clean the text (the same as above)
txt = re.sub(r'(\n\[)(.*?)(\[]\n)', '\n', txt, flags=re.DOTALL)

print(txt)

Read the file into a string,

extract = '''Hi this is my config file.
Please dont delete it

[homes]
  browseable                     = No
  comment                        = Your Home
  create mode                    = 0640
  csc policy                     = disable
  directory mask                 = 0750
  public                         = No
  writeable                      = Yes

[proj]
  browseable                     = Yes
  comment                        = Project directories
  csc policy                     = disable
  path                           = /proj
  public                         = No
  writeable                      = Yes

[]

This last second line.
End of the line.
'''.split('\n[')[0][:-1]

will give you,

Hi this is my config file.
Please dont delete it

.split('\n[') splits the string by the occurrence of '\n[' set of characters and [0] selects the upper description lines.

with open("smb", "r") as f: 
     extract = f.read()
     tail = extract.split(']\n')
     extract = extract.split('\n[')[0][:-1]+[tail[len(tail)-1]

will read and output,

Hi this is my config file.
Please dont delete it
This last second line.
End of the line.

Since you tagged pandas, let's try that:

df = pd.read_csv('smb', sep='----', header=None)

# mark rows starts with `[`
s = df[0].str.startswith('[')

# drop the lines between `[`
df = df.drop(np.arange(s.idxmax(),s[::-1].idxmax()+1))

# write to file if needed
df.to_csv('clean.txt', header=None, index=None)

Output (df):

                             0
0   Hi this is my config file.
1        Please dont delete it
18      This last second line.
19            End of the line.

You can iterate over file lines and collect them into some list unless reach line wrapped into brackets, then concatenate collected lines back:

with open("smb", "r") as f:
    result = []
    for line in f:
        if line.startswith("[") and line.endswith("]"):
            break
        result.append(line)
    result = "\n".join(result)
    print(result)

If I understand you correctly, you want everything before the first [ and after the last ]. If it is not the case please let me know and I will change my answer.

with open("smb", "r") as f: 
    s = f.read()
    head = s[:s.find('[')]
    tail = s[s.rfind(']') + 1:]
    return head.strip("\n") + "\n" + tail.strip("\n") # removing \n

This will give you the desire output.

Another option is to first match the the square brackets like [homes], then match all lines that do not only contain [] as that is the end marker.

You could get the match without using (?s) or using re.DOTALL to prevent unnecessary backtracking and replace the match with an empty string.

^\s*\[[^][]*\](?:\r?\n(?![^\S\r\n]*\[]$).*)*\r?\n[^\S\r\n]*\[]$\s*

Explanation

  • ^ Start of line
  • \s* Match 0+ whitepace chars
  • \[[^][]*\]
  • (?: Non capture group
    • \r?\n Match a newline
    • (?! Negative lookahead, assert what is on the right is not
      • [^\S\r\n]*\[]$ match 0+ times a whitespace char except newlines and match []
    • ) Close non capture group
    • .* Match 0+ times any char except a newline
  • )* Close non capture group and repeat 0+ times
  • \r?\n Match a newline
  • [^\S\r\n]* Match 0+ whitespace chars without a newline
  • \[]$ Match [] and assert the end of the line
  • \s* Match 0+ whitespace characters

Regex demo | Python demo

Code example

import re

regex = r"^\s*\[[^][]*\](?:\r?\n(?![^\S\r\n]*\[]$).*)*\r?\n[^\S\r\n]*\[]$\s*"

with open("smb", "r") as file:
    data = file.read()
    result = re.sub(regex, "", data, 0, re.MULTILINE)
    print(result)

Output

Hi this is my config file.
Please dont delete it
This last second line.
End of the line.

On Regex101 you can test this:

(^\W)+?\[[\w\W]+?\[\](\W)+?(\w)

In code its like

import re ------------------------------------------------------------↧-string where to replace-- result = re.sub(r"(^\W)+?\[[\w\W]+?\[\](\W)+?(\w)", "", input_string, 0, re.MULTILINE) ----------------------↑-this is the regex------------↑-substitution string-------------

Cheers

Since you've tagged pandas and stipulate that the text comes before & after the square brackets we can use str.contains and use a boolean to filter out the rows that fall in between the first & last square bracket.

df = pd.read_csv(your_file,sep='\t',header=None)

idx = df[df[0].str.contains('\[')].index

df1 = df.loc[~df.index.isin(range(idx[0],idx[-1] + 1))]

                             0
0   Hi this is my config file.
1        Please dont delete it
18      This last second line.
19            End of the line.

You've got the indexing wrong. Apart from that, the code seems fine.

Try:

start=0
targ = ""
end=0
with open("smb", "r") as file:
    for line in file: 
        try:  
            if start==0:
                start = line.index("[")
        except:
            start = start
        try:  
            end = line.index("]")
        except:
            end = end
        targ = targ+line

targ = targ[0:start-1]+targ[end+1:]

This should work. Let me know if anything goes wrong. :)

Using Pandas :

df = pd.read_csv('smb.txt', sep='----', header=None, engine='python',names=["text"])

res = df.loc[~df.text.str.contains("=|\[.*\]")]
print(res)
text
0   Hi this is my config file.
1   Please dont delete it
18  This last second line.
19  End of the line.

Explanation : Exclude rows that contain either = or contain a starting bracket ([) that may or may not be followed by characters(.*) and have a closing bracket (]``). the backslash (```) tells python not to treat the brackets as special characters

With Python only, the same regex pattern is used, with an extra line to take care of empty entries:

import re
with open('smb.txt') as myfile:
    content = myfile.readlines()
    pattern = re.compile("=|\[.*\]")
    res = [ent.strip() for ent in content if not pattern.search(ent) ]
    res = [ent for ent in res if ent != ""]
    print(res)
['Hi this is my config file.',
 'Please dont delete it',
 'This last second line.', 
 'End of the line.']

Here is probably one of the cleanest ways you can do it.

import re
from pathlib import Path
res = '\n'.join(re.findall(r'^\w.*', Path('smb').read_text(), flags=re.M))

Explanation:

Path creates a Path object to the file. Path.read_text() opens the file reads the text and closes the file. The file contents are passed to re.findall which uses the re.M flag to look at each line in the file to validate again the pattern '^\w.*' which will only take lines that begin with a word character. This eliminates lines that begin with white-space or brackets.

Try r"(?s)\s*\[[^\[\]]*\](?:(?:(?!\[[^\[\]]*\]).)+\[[^\[\]]*\])*\s*"
Replace r"\n"

demo

Related