Remove specific items (0 values and values multiplied by *0) from a large text file and write it to a new text file using Python

Viewed 172

I am a basic Python user and I have searched in multiple platforms how to delete from a large text file specific values but I haven't found anything similar to what I want to do . I have a large file (out.txt) and I want to remove all the 0 values and all values multiplied by 0 (75*0) in the large data file. After removing all those values I want to write it in a new text file (out2.txt). Suggestions please. Thanks!

I have tried this code;

content = open('out.txt', 'r').readlines()

content_set = set(content)

cleandata = open('clean.txt', 'w')

for line in content_set:

    cleandata.remove(0)

I keep getting this error:

cleandata.remove(0)

AttributeError: '_io.TextIOWrapper' object has no attribute 'remove'

DATA FILE out.txt

75*0 78.8502 45.9301 13358*0 10.7678 0 23.9901 43.8503 77*0 1.3757 36.9888 15.0398 76*0 8.19519 0 4.11938 21.4933 23.832 76*0 34.7566

  15.5595 21.0239 0 47.1607 76*0 14.9065 52.916 51.7825 13358*0 62.4689 22.8217 15.68 77*0 12.8943 0 32.1276 14.1273 76*0 39.6095

  70.8503 72.8765 45.7607 76*0 12.5657 72.7567 58.0161 30.9 76*0 19.5879 648.696 111.501 13358*0 17.36 18.0555 85.0358 77*0 4.62265

  55.7498 61.2049 76*0 762.354 8.34207 23.2367 16.0517 76*0 405.637 20.1265 8.17844 16.4698 76*0 107.228 35.1968 38.4117 13358*0
3 Answers

Try this:

with open('out.txt') as f:
    s=f.read()

s=' '.join([i for i in s.split(' ') if i!='0' and '*0' not in i])

with open('out2.txt', 'w') as f:
    f.write(s)

Output:

78.8502 45.9301 10.7678 23.9901 43.8503 1.3757 36.9888 15.0398 8.19519 4.11938 21.4933 23.832 34.7566

15.5595 21.0239 47.1607 14.9065 52.916 51.7825 62.4689 22.8217 15.68 12.8943 32.1276 14.1273 39.6095

70.8503 72.8765 45.7607 12.5657 72.7567 58.0161 30.9 19.5879 648.696 111.501 17.36 18.0555 85.0358 4.62265

55.7498 61.2049 762.354 8.34207 23.2367 16.0517 405.637 20.1265 8.17844 16.4698 107.228 35.1968 38.4117

This is should work:

content = open('out.txt', 'r').readlines()

cleandata = []
for line in content:
    line = {i:None for i in line.replace("\n", "").split()}
    for value in line.copy():
        if value == "0" or value.endswith("*0"):
            line.pop(value)
    cleandata.append(" ".join(line) + "\n")

open('clean.txt', 'w').writelines(cleandata)
content = open("out.txt").read()
segments = content.split()
for segment in range(len(segments)):
    if segments[segment]=="0" or segments[segment].endswith("*0"):
        del segments[segment]
clean = open("clean.txt", "w")
clean.write(" ".join(segments))
clean.close()

What this does is take all of the content of out.txt and split() it on all whitespace (no argument means all whitespace). Then it loops over each segment and on each segment, checking if the segment is 0 or contains *0, and if it has either, deletes the segment from segments. At the end, it creates clean.txt, writes all of the segments with spaces separating them, and then closes clean.txt.

The only problem with this that I noticed is that when it writes to clean.txt, they are separated by spaces instead of their original whitespace. One way to fix this is to store the whitespace after each number and when it contains 0 or includes *0, destroy the segment and it's associated whitespace.

Try it and tell me in the comments if it works!

Related