How to extract (First match)text between two words

Viewed 473

I have a file having the following structure

destination list

move from station d-435-435 to point place1
move from station d-435-435 to point place2
move from mainpoint

I want to extract the word "d-435-435"(Only the first match, this need not be same value always) in between the words "from station" and "to point"

How can I achieve this?

What I have tried so far?

id=$(sed 's/.*from station \(.*\) to.*/\1/' input.txt)

But this returns the following value: destination list d-435-435 move from mainpoint

5 Answers

1st solution: With your shown samples, please try following GNU awk code. Using match function of awk program here to match regex rom station\s+\S+\s+to point to get requested value by OP then removing from station\s+ and \s+to point from matched value and printing required value.

awk '
match($0,/from station\s+\S+\s+to point/){
  val=substr($0,RSTART,RLENGTH)
  gsub(/from station\s+|\s+to point/,"",val)
  print val
  exit
}
' Input_file


2nd solution: Using GNU grep please try following. Using -oP option to print matched portion and enabling PCRE regex respectively here. Then in main grep program matching string from station followed by space(s) then using \K option will make sure matched part before \K is forgotten(since e don't need this in output), Then matching \S+(non space values) followed by space(s) to point string(using positive look ahead here to make sure it only checks its present or not but doesn't print that).

grep -oP -m1 'from station\s+\K\S+(?=\s+to point)' Input_file

If GNU sed is available, how about:

id=$(sed -nE '0,/from station.*to/ s/.*from station (.*) to.*/\1/p' input.txt)
  • The -n option suppress the print unless the substitution succeeds.
  • The condition 0,/pattern/ is a flip-flop operator and it returns false after the pattern match succeeds. The 0 address is a GNU sed extension which makes the 1st line to match against the pattern.

With awk you can write the before and after conditions of field $4, where d-435-435 is, and then print this field only the first match and exit with exit after print statement:

awk '$2=="from" && $3=="station" && $5=="to" && $6=="point" {print $4; exit}' file
d-435-435

or using GNU awk for the 3rd arg to match():

awk 'match($0,/from station\s+(.*)\s+to point/,a){print a[1];exit}' file
d-435-435
  • The regexp contains a parenthesis, so the integer-indexed element of array a[1] contain the portion of string between from station followed by space(s) \s+ and space(s) \s+ followed byto point.

This might work for you (GNU sed):

sed -nE '/.*station (\S+) to point.*/{s//\1/;H;x;/\n(\S+)\n.*\1/{s/\n\S+$//;x;d};x;p}' file

Turn off implicit printing and on extended regexps command line options -nE.

If a line matches the required criteria, extract the required string, append a copy to the hold space, check if the match has already been seen and if not print it. If the match has been seen, remove it from the hold space.

Otherwise, do not print anything.

This should work in any sed:

sed -e '/.*from station \([^ ]*\) to .*/!d' -e 's//\1/' -e q file
Related