Python URL join beginning with ////

Viewed 48

In HTML, we find URLs like ////assets.example.com/some/ressource.

A web browser on page https://www.example.com/original/page will construct the following URL from that:

https://assets.example.com/some/ressource

But when using the URLjoin method in python

from urllib import parse
parse.urljoin("https://www.example.com/original/page", "////assets.example.com/some/ressource")

we get

https://www.example.com//assets.example.com/some/ressource

Why do web browsers behave different than the URLjoin method here? Who is right here?

1 Answers

On the question of "who is right?" The major difference in this is that Python is stricter about what you're giving it than browsers are when it comes to respecting URLs per RFC-9386.

But note that urllib.parse.urljoin does something more specific than just "parse URLs for correctness" because its purpose is to construct an absolute URL from what you give it. When you give it ////..., it ends up with a parsed URI that doesn't have a netloc (or host, as you will) so behaviorally, it has to fall-back on using the base URL to see if it has a domain to construct an absolute URL.

And if you wanted to achieve the same thing as how a browser would do it, use the 'expected' format of just //host/path rather than more than two /, like so...

new_url = parse.urljoin(
    "https://www.example.com/original/page", 
    "//assets.example.com/some/resource"
)
print(new_url)

# output
# https://assets.example.com/some/resource

Read further if you want to dive into the internals of how Python parses URIs for urljoin().


When you call urljoin, it starts off by parsing the URLs you pass it using urllib.parse.urlparse:

# urllib.parse.urljoin
def urljoin(base, url, allow_fragments=True):
    """Join a base URL and a possibly relative URL to form an absolute
    interpretation of the latter."""
    if not base:
        return url
    if not url:
        return base

    base, url, _coerce_result = _coerce_args(base, url)
    bscheme, bnetloc, bpath, bparams, bquery, bfragment = \
            urlparse(base, '', allow_fragments)
    scheme, netloc, path, params, query, fragment = \
            urlparse(url, bscheme, allow_fragments)
    # ...

We can dive into how URLs are parsed, and inside urllib.parse.urlparse, it executes the following:

# urllib.parse.urlparse
def urlparse(url, scheme='', allow_fragments=True):
    url, scheme, _coerce_result = _coerce_args(url, scheme)
    splitresult = urlsplit(url, scheme, allow_fragments)
    scheme, netloc, url, query, fragment = splitresult
    
    if url[:2] == '//':
        netloc, url = _splitnetloc(url, 2)
        if (('[' in netloc and ']' not in netloc) or
                (']' in netloc and '[' not in netloc)):
            raise ValueError("Invalid IPv6 URL")

With _splitnetloc being:

# urllib.parse._splitnetloc
def _splitnetloc(url, start=0):
    delim = len(url)   # position of end of domain part of url, default is end
    for c in '/?#':    # look for delimiters; the order is NOT important
        wdelim = url.find(c, start)        # find first of this delim
        if wdelim >= 0:                    # if found
            delim = min(delim, wdelim)     # use earliest delim position
    return url[start:delim], url[delim:]   # return (domain, rest)

Take all this code together, and when you pass it something of the form ////name/path, the URL parser finds // to start, and then tries to find the netloc with the subsequent string //name/path, which _splitnetloc determines would be ('', '//name/path') because it searches for the first / possible to denote the path of the URL.

So see that your two assets.example.com URLs would be parsed differently based on this logic:

In [1]: four = "////assets.example.com/some/ressource"
In [2]: two = "//assets.example.com/some/ressource"
In [3]: pfs = urllib.parse.urlparse(four, "https")
In [4]: pts = urllib.parse.urlparse(two, "https")

In [5]: print(pfs)
ParseResult(scheme='https', netloc='', path='//assets.example.com/some/ressource', params='', query='', fragment='')

In [6]: print(pts)
ParseResult(scheme='https', netloc='assets.example.com', path='/some/ressource', params='', query='', fragment='')

The absence of the netloc is important because, going way back to the beginning when I talked about urljoin, the subsequent code after it performs the urlparse of the two input URLs, is this:

if scheme in uses_netloc:
    # this is the netloc from the second argument URL
    if netloc:
        return _coerce_result(urlunparse((scheme, netloc, path,
                                          params, query, fragment)))
    netloc = bnetloc

And so, with ////assets.example.com/..., the parsed URL doesn't have a netloc and so it ends up that you use the base URL domain (netloc = bnetloc line).

Whereas with //assets.example.com/..., it does have a netloc so the code just immediately returns urlunparse(...) on the parsed URL, resulting in a url like https://assets.example.com/....

Related