How can I parse info with Logstash?

Viewed 64

I have an input:

May 16 12:45:47 host-dev1 kernel: [  162.648366] wireguard: wg0: Sending keepalive packet to peer 2 (171.12.198.123:51079)

I want to parse the info as: TIMESTAMP "Sending keepalive packet to peer 2" IP:PORT

For the middle sentence I want to parse whatever is after wg0: until the first parenthesis of the port. This sentence can change to "Sending handshake initiation to peer 10" for example.

I've done

filter {
    grok {
        match => { "message" => "%{SYSLOGBASE:timestap} %{GREEDYDATA:action} %{IP:peerip}:%{NUMBER:port}" }
    }
}

I need to change GREEDYDATA to something that will specifically parse the mentioned boundaries

1 Answers

Give this a try:

%{SYSLOGBASE:timestamp} \[ ?%{NUMBER:TIMESTAMP} ?\]( %{WORD}:)* %{GREEDYDATA:action} \(%{IP:peerip}:%{NUMBER:port}

Here's a breakdown:

%{SYSLOGBASE:timestamp}       - the syslog prefix
\[ ?%{NUMBER:TIMESTAMP} ?\]   - the application timestamp
( %{WORD}:)*                  - any words followed by a colon, like
                                'wg0:', zero or more times
%{GREEDYDATA:action}          - any characters ('DATA' would also work)
\(%{IP:peerip}:%{NUMBER:port} - a literal '(' followed by IP and port

The important thing in making GREEDYDATA / DATA work here is that the the boundaries (%{WORD}: and \() are properly defined.

You may need to vary the boundary definitions depending on what other log messages look like (specifically, whether you can rely on the colon and parenthesis at the boundaries).

It may be helpful to use a named capture group, depending on whether existing grok patterns cover your other message formats, like : (?<notColons>[^:]*) \( to specify "a colon, then a space, then any number of non-colon characters, then a space, then an open bracket".

Related