Regex groups for cloudinary url

Viewed 223

I'm trying to capture different parts of a url while ignoring parts that sometimes comes up.

I've tried using and extending the regex found here with little luck. https://gist.github.com/ahmadawais/9813c44b7e51c2c3540d2165d6c6cc65

Take the example

https://res.cloudinary.com/test-site/image/upload/v1619174590/folder/path/cjtdn73cleqagpy4fqza.jpg

https://res.cloudinary.com/test-site/image/upload/ar_1:1,c_fill,f_auto,g_auto,w_700/v1619174590/folder/path/cjtdn73cleqagpy4fqza.jpg

https://res.cloudinary.com/test-site/image/facebook/fb_id

res.cloudinary.com : host

test-site : cloudname

upload/facebook: resource_type

v1619174590/rg/collective/media/cjtdn73cleqagpy4fqza.jpg: id

I need to ignore everything between /upload/ and /v, I've accomplished this using //upload/.*?\b(?=v1)/ , but it doesn't account for if the resource type is facebook and there is no /v123

2 Answers

I am assuming that your question is specific to Cloudinary URL format. If that is correct, the URL format will follow this pattern:

  1. Protocol (http or https)
  2. Domain (res.cloudinary.com)
  3. Cloud name / Sub-Account name
  4. Resource type (image, video or raw)
  5. Visibility (upload, authenticated)
  6. Transformation (or chained transformations)
  7. Version number
  8. Path to your resource also called as public-id in Cloudinary terms
  9. Extension (note that extension is not considered part of public-id in Cloudinary)

In your example URL https://res.cloudinary.com/test-site/image/upload/ar_1:1,c_fill,f_auto,g_auto,w_700/v1619174590/folder/path/cjtdn73cleqagpy4fqza.jpg, this would map as follows:

  1. https - protocol
  2. res.cloudinary.com - domain
  3. test-site - cloud name
  4. image - resource type
  5. upload - visibility (ie a public asset)
  6. ar_1:1,c_fill,f_auto,g_auto,w_700 - transformation
  7. v1619174590 - version number
  8. folder/path/cjtdn73cleqagpy4fqza
  9. jpg - extension. Without f_auto, the result would have been a JPG image.

Using this logic, the regex to catch most URLs would be as follows:

(https?)\:\/\/(res.cloudinary.com)\/([^/]+)\/(image|video|raw)\/(upload|authenticated)\/(.*)\/(v[0-9]+)\/(.+)(?:\.[a-z]{3})?

You can use

https?:\/\/(?<host>[^\/]+)\/(?<cloudname>[^\/]+)\/[^\/]+\/(?<resource_type>[^\/]+)(?:\/[^\/,]*,[^\/]*)?\/(?<id>.*)
https?:\/\/([^\/]+)\/([^\/]+)\/[^\/]+\/([^\/]+)(?:\/[^\/,]*,[^\/]*)?\/(.*)

The first regex is compliant with the ECMAScript 2018+ standard that supports named capturing groups, and the second one just contains regular, numbered capturing groups.

See the regex demo.

Details

  • https?:\/\/ - https:// or http://
  • ([^\/]+) - Group 1 (host): one or more chars other than / - \/ - a / char
  • ([^\/]+) - Group 2 (cloud name): one or more chars other than /
  • \/[^\/]+\/ - /, any one or more chars other than / and a /
  • ([^\/]+) - Group 3 (resource type): one or more chars other than /
  • (?:\/[^\/,]*,[^\/]*)? - an optional sequence of
    • \/ - a / char
    • [^\/,]* - zero or more chars other than / and ,
    • , - a comma
    • [^\/]* - zero or more chars other than /
  • \/ - a / char
  • (.*) - Group 4 (id): the rest of the string.
Related