extract filename out of raw data using regex

Viewed 418

This is my raw data:

h24-71-249-14.ca.shawcable.net - - [07/Mar/2004:22:29:13 - 0800] "GET /icons/gnu-head-tiny.jpg HTTP/1.1" 200 3049

h24-71-249-14.ca.shawcable.net - - [07/Mar/2004:22:29:13 - 0800] "GET /icons/gnu-head-tiny HTTP/1.1" 200 3049

I want to be able to extract a file's name from the URI (if there is any, if there is not - ignore). The file can be any filetype (jpg, png, txt, etc.)

This is what I have so far:

(\"+)(.*?)(\.\w{1,3})

I know it is probably not a good idea to start my string from ", and it is probably the reason for my problem, so I'd like to get some help to fix my regex.

thank you!

2 Answers

You can start the pattern from ", but you don't have to escape, repeat and capture it.

If you want the extension together with the filename, you can use a single capturing group.

You might use:

"GET \S+\/(\S+\.\w{1,3})\b

Explanation

  • "GET Match literally
  • \S+/ Match 1+ non whitespace chars and then match the last /
  • (\S+\.\w{1,3}) Capture group 1, match 1+ non whitespace chars, a dot and 1-3 word chars
  • \b A word boundary

Regex demo

There is no language tagged, but for example using Javascript

const regex = /"GET \S+\/(\S+\.\w{1,3})\b/;
[
  "h24-71-249-14.ca.shawcable.net - - [07/Mar/2004:22:29:13 - 0800] \"GET /icons/gnu-head-tiny.jpg HTTP/1.1\" 200 3049",
  "h24-71-249-14.ca.shawcable.net - - [07/Mar/2004:22:29:13 - 0800] \"GET /icons/gnu-head-tiny HTTP/1.1\" 200 3049"
].forEach(s => {
  let m = s.match(regex);
  if (m) console.log(m[1]);
})


When \K is supported, you can get the match only. According to the comments, this pattern gets the specific desired match:

\w{1,5} \S+\/\K\S+\.\w{3}\b

Explanation

  • \w{1,5} Match 1-5 word chars and a space
  • \S+\/ Match 1+ non whitespace chars and then the last /
  • \K Reset the match buffer (forget what is matched until now)
  • \S+ Match 1+ non whitespace chars
  • \.\w{3} Match a dot and 1-3 word characters
  • \b A word boundary

Regex demo

Here are two options:

First

If you want what's between the GET and HTTP, this will do it:

| rex field=_raw "GET\s+(?<fname>\S+)\s+HTTP"

Start at the string literal GET, go one (or more) whitespaces, then put everything that's not a whitespace character (up until a whitespace sequence that ends in the string literal HTTP) into the new field fname.

Functionally, you can leave off the \s+HTTP from the regex, but for fullness' sake, you may want to choose to leave it in there.

Second

If you only want the ending filename, this is it:

| rex field=_raw "(?<fname>[\.\-\w]+)\s+HTTP"

This will match all instances of ., -, and any word character (\w) as many times as they are found before a sequence of whitespace characters (\s+) followed by the string literal HTTP into the new field fname.

Or, optionally (though more steps to find the match, it might be better in your case):

| rex field=_raw "(?<fname>[^\/]+)\s+HTTP"

This one will match anything that is not a front slash (/) up to the series of whitespaces followed by HTTP all into the new field fname.

Related