Matching new line in TCL - regexp

Viewed 178

I have a variable

set a "--------------------------------------------------------------------------------
       Proto     Source Address                         Pkt-Cnt    Start
                 Destination Address                    Byte-Cnt
       --------------------------------------------------------------------------------

       UDP       150.1.1.2                              25         05/24/2021 07:07:29
                 150.2.1.2                              1150      

      --------------------------------------------------------------------------------"

i need to match all the values after the word UDP. i tried this one to match the first line and it works good. But i could not get the values "150.2.1.2" and "1150" - Any help is much appreciated

    regexp "UDP + (\[\[:graph:]]+) + (\[\[:graph:]]+) + (\[\[:graph:]]+) +(\[\[:graph:]]+)"  $a match data1 data2 data3 data4
2 Answers

There's a few ways to do it, but I think this RE is what you ought to use:

{UDP\s*([\d.]+)\s*(\d+)\s*([\w/]+ [\w:]+)\s*([\d.]+)\s*(\d+)}

It's enclosed in braces because otherwise there's a lot of extra backslashes!

The essential pieces:

  1. UDP — Marker text
  2. \s* — Whitespace (spaces, tabs, etc)
  3. ([\d.]+) — Captured digits and dots (the source address)
  4. \s* — Whitespace
  5. (\d+) — Captured digits (the packet count)
  6. \s* — Whitespace
  7. ([\w/]+ [\w:]+) — Captured start timestamp (with only a single space in the middle)
  8. \s* — Whitespace (this includes the newline; newlines are whitespace by default)
  9. ([\d.]+) — Captured digits and dots (the destination address)
  10. \s* — Whitespace
  11. (\d+) — Captured digits (the byte count)

In use:

regexp {UDP\s*([\d.]+)\s*(\d+)\s*([\w/]+ [\w:]+)\s*([\d.]+)\s*(\d+)} $a -> source packetCount start destination byteCount

You can follow the similar logic and add two more capturing groups separated with \s+ (one or more whitespaces):

regexp {UDP\s+(\S+)\s+(\S+)\s+(\S+)\s+(\S+)\s+(\S+)\s+(\S+)}  $a match data1 data2 data3 data4 data5 data6

Note that \S+ matches one or more non-whitespace chars and \s+ matches one or more whitespace chars.

See the online Tcl demo with UDP declared as a variable:

set a "--------------------------------------------------------------------------------
       Proto     Source Address                         Pkt-Cnt    Start
                 Destination Address                    Byte-Cnt
       --------------------------------------------------------------------------------

       UDP       150.1.1.2                              25         05/24/2021 07:07:29
                 150.2.1.2                              1150      

      --------------------------------------------------------------------------------"

set b "UDP"
regexp "$b\\s+(\\S+)\\s+(\\S+)\\s+(\\S+)\\s+(\\S+)\\s+(\\S+)\\s+(\\S+)" $a match data1 data2 data3 data4 data5 data6
puts "$data1, $data2, $data3, $data4, $data5, $data6"
# 150.1.1.2, 25, 05/24/2021, 07:07:29, 150.2.1.2, 1150
Related