How to extract only the first instance of a number of lines between two strings in bash?

Viewed 556

My file is:

abc
123
xyz
abc
675
xyz

And I want to extract:

abc
123
xyz

(123 could be anything, the point is I want the first occurrence)

I tried using this:

sed -n '/abc/,/xyz/p' filename

but this is giving me all the instances. How could I get just the first one?

6 Answers

Could you please try following, written and tested with shown samples.

awk '/abc/{found=1} found; /xyz/ && found{exit}'  Input_file

OR as per Ed sir's comment for better efficiency try following.

awk '/abc/{found=1} found{print; if (/xyz/) exit}'  Input_file

Explanation: Adding detailed explanation for above.

awk '               ##Starting awk program from here.
/abc/{              ##checking condition if a line has abc in it then do following.
  found=1           ##Setting found here.
}
found;              ##Checking condition if found is SET then print that line.
/xyz/ && found{     ##Checking if xyz found in line and found is SET then do following.
  exit              ##exit program from here.
}
'  Input_file       ##Mentioning Input_file name here.

If you don't mind Perl:

perl -ne 'm?abc?..m?xyz? and print' file

will print only the first block that matches. The delimiter for the matches must be the ? character.

Using sed you can do:

sed -n '/abc/,/xyz/p; /xyz/q' filename

q will quit after the "xyz" pattern is reached.

Match the Terminal Condition Twice

Regardless of the language, the most common technique for line-oriented processing is to print lines within a given range and then use a second command to exit the loop when your terminal condition is reached. This will be true for common patterns in sed, awk, ruby, and perl, although there are certainly other techniques that can be performed using multi-line matches (not supported in sed without using the hold space). For example, you might use a non-greedy, multi-line regular expression pattern such as /^abc\n.*?\nxyz$/m.

To illustrate the line-oriented approach you want a little more verbosely, consider this Ruby one-liner where $_ holds the current input line. From the shell:

$ ruby -ne 'puts $_ if /^abc$/ .. /^xyz$/; exit if /^xyz/' filename 
abc
123
xyz

The equivalent in sed is:

$ sed -n '/^abc$/,/^xyz$/p; /^xyz$/q' filename
abc
123
xyz

All you were missing was a quit or exit command attached to the second match against the first instance of xyz.

This has already been well and sufficiently answered and hashed out by better minds than me, but

  1. since you explicitly used sed, and
  2. for some variety of approach that handles the requested conditions...
sed -n '/abc/,/xyz/{ p; /xyz/q; }' filename
  • This only looks at the range in question, so won't print or quit on xyz with no opening abc ahead of it
  • it prints all the records in the range
  • it exits on the first xyz it sees AFTER an abc, so will dependably exit
  • if the end doesn't have an xyz it will just print to EOF.

You might refine the pattern if you want to make sure of exact matching, such as

sed -n '/^abc$/,/^xyz$/{ p; /^xyz$/q; }' filename

This prevents near-matches from confusing the logic, but is (intentionally) unforgiving of botched sentinel strings.

This might work for you (GNU sed):

 sed '/abc/!d;:a;n;/xyz/!ba;q' file

If it is not a line containing abc delete it.

Otherwise, print it and fetch the next.

If that line is not xyz, repeat.

Otherwise, quit.

N.B. The option -n is not set and so the last line will be printed before termination.

This will print until the end of the file or the string xyz is encountered.

If xyz must be present, use:

sed -n '/abc/!d;:a;N;/xyz/!ba;p;q' file
Related