Remove line from file that contains string more than once

Viewed 55

I need to remove all the lines inside a file that contains a certain string more than once, for example if my file is like this:

This is a test toRemove first line

This is a test toRemove second line toRemove

Should produce a file like with only the first line

This is a test toRemove first line

I am trying to do this on linux from command line and I tried to use grep or sed like this

grep -d "toRemove.*toRemove" myFile > myOtherFile

sed '/\toRemove.*toRemove/!d' myFile > myOtherFile

But nothing seems to work. Does anybody know how to obtain this?

2 Answers

You can use

sed '/toRemove.*toRemove/d' myFile > myOtherFile
grep -v "toRemove.*toRemove" myFile > myOtherFile

sed: Note that \t matches a TAB char, and !d removes lines that do not match the pattern. So, you need to remove the \ before t and remove ! before the d.

grep: You should have used the -v option that reverses the result of the regex check (it will output all lines that do NOT match the pattern).

See the online demo:

s='This is a test toRemove first line
This is a test toRemove second line toRemove'
sed '/toRemove.*toRemove/d' <<< "$s"
# => This is a test toRemove first line
grep -v 'toRemove.*toRemove' <<< "$s"
# => This is a test toRemove first line

This might work for you (GNU sed):

sed '/\<\(toRemove\)\>.*\<\1\>/d' file

This will delete a line that has 2 or more occurrences of the word toRemove.

To delete any line that contains the word toRemove beyond the first such line:

sed '/\<toRemove\>/{x;/./{x;d};x;h}' file
Related