Using regex to parse data between delimeter and ending at a specified substring

Viewed 31

I'm trying to parse out the names from a bunch of semi-unpredictable strings. More specifically, I'm using ruby, but I don't think that should matter much. This is a contrived example but some example strings are:

Eagles vs Bears
NFL Matchup: Philadelphia Eagles VS Chicago Bears TUNE IN
NFL Matchup: Philadelphia Eagles VS Chicago Bears - TUNE IN
Philadelphia Eagles vs Chicago Bears - NFL Match
Phil.Eagles vs Chic.Bears
3agles vs B3ars

The regex I've come up with is

([0-9A-Z .]*) vs ([0-9A-Z .]*)(?:[ -:]*tune)?/i

but in the case of "NFL Matchup: Philadelphia Eagles VS Chicago Bears TUNE IN" I'm receiving Chicago Bears TUNE as the second match. I'm trying to remove "tune in" so it's in it's own group.

I thought that by adding (?:[ -:]*tune)? it would separate the ending portion of the expression the same way that having vs in the middle was able to, but that doesnt seem to be the case. If I remove the ? at the end, it matches correctly for the above example, but it no longer matches for Eagles vs Bears

If anyone could help me, I would greatly appreciate it if you could breakdown your regex piece by piece.

2 Answers

You can capture the second group up to a -, : or tune preceded with zero or more whitespaces or till end of the line while making the second group pattern lazy:

([\w .]*) vs ([\w .]*?)(?=\s*(?:[:-]|tune|$))

See the regex demo.

Details:

  • ([\w .]*) - Group 1: zero or more word, space or . chars as many as possible
  • vs - a vs string
  • ([\w .]*?) - Group 2: zero or more word, space or . chars as few as possible
  • (?=\s*(?:[:-]|tune|$)) - a positive lookahead that requires the following pattern to appear immediately to the right of the current location:
    • \s* - zero or more whitespaces
    • (?:[:-]|tune|$) - : or -, tune or end of a line.

You can use the following regular expression which I have expressed in free-spacing mode to make it self-documenting (search for "Free-Spacing Mode" at the link).

 rgx = /
       (?: |\A)              # match space or beginning of string
       (?<team1>             # begin capture group team1
         (?<team>            # begin capture group team
           (?<word>          # begin capture group word
             (?:\p{Lu}|\d)   # word begins with an uppercase letter or digit
             (?:\p{Ll}|\d)+  # ...followed by 1+ lowercase letters or digits
           )                 # end capture group word
           (?:               # begin non-capture group
             [ .]            # match a space or period
             \g<word>        # match another word
           )*                # end non-capture group and execute 1+ times
         )                   # end capture group team
       )                     # end capture group team1
       [ ]+                  # match one or more spaces
       (?:VS|vs)             # match literal
       [ ]+                  # match one or more spaces
       (?<team2>             # begin capture group team2
         \g<team>            # match the second team name
       )                     # end capture group team2
       (?:                   # begin non-capture group
         [ ]                 # match a space
         (?:                 # begin non-capture group
           (?:-[ ])?         # optionally match literal
           TUNE[ ]IN         # match literal
           |                 # or
           -[ ]NFL[ ]Match   # match literal
         )                   # end inner capture group 
       )?                    # end outer non-capture group and make it optional
       \z                    # match end of string
       /x                    # free-spacing regex definition mode     
examples = [
  "Eagles vs Bears",
  "NFL Matchup: Philadelphia Eagles VS Chicago Bears TUNE IN",
  "NFL Matchup: Philadelphia Eagles VS Chicago Bears - TUNE IN",
  "Philadelphia Eagles vs Chicago Bears - NFL Match",
  "Phil.Eagles vs Chic.Bears",
  "3agles vs B3ars"
]
examples.map do |s|
  m = s.match(rgx)
  [m[:team1], m[:team2]]
end
  #=> [["Eagles", "Bears"],
  #    ["Philadelphia Eagles", "Chicago Bears"],
  #    ["Philadelphia Eagles", "Chicago Bears"],
  #    ["Philadelphia Eagles", "Chicago Bears"],
  #    ["Phil.Eagles", "Chic.Bears"],
  #    ["3agles", "B3ars"]]

See Regexp#match and MatchData#[].

Note that \g<word> and \g<team> effectively copy the code contained in the capture groups word and team, respectively. These are called "Subexpression Calls". For additional information search for that term at Regexp. There are two advantages to using subroutine calls: less code is needed and the opportunities for coding errors is reduced.

Related