Python 3.10.2
I have a URL that generally appears as follows, with some slight variation (http/https, www. prefix sometimes, #params at the end to indicate things like referrer or device being displayed on, etc).
https://madeupdomain.net/u/Hypothetical_Username/Some-Random-Page-Name
The forms of the URL I'm generally encountering are either:
https://madeupdomain.net/u/Hypothetical_Username/
or
https://madeupdomain.net/u/Hypothetical_Username/Some-Random-Page-Name
What I'm interested in doing with the URL:
- Get the
Hypothetical_Usernamepart - Find out if the URL stops at the username or if there is another
/pathafter it
I've been using user = url.split('/')[4] to get the username portion of the URL. Since the URL always includes the username and the URL is usually consistent (for now), I can rely on this split getting the element I want. If the URL changes a little bit in the future, I know this'll bone me.
However, the rest of the path is optional.
If I just use url.split('/')[5], python throws an error as soon as it encounters a URL where split doesn't have a [5]th element.
So I tired to "test" for it with an if statement and it still complains and throws the error IndexError: list index out of range.
if url.split('/')[5]:
continue
When I print out the list, it'll look like either of the following. As you can see, there are 5 elements in the first and six in the second.
['https:', '', 'madeupdomain.net', 'u', 'Hypothetical_Username']
['https:', '', 'madeupdomain.net', 'u', 'Hypothetical_Username', 'Some-Random-Page-Name']
So, I tried running len(url.split('/')) on every iteration, to see how many elemnts each list has and it always says 6 - whether it is the first or second example, above.
So, I'm kind of at a loss here as to a very simple and clean way to do what I want to do. I know there are url parsing libraries, but that seems like overkill for what I want to do (get the username, then find out if there is a path name beyond that and decide what to do with the URL once I know).
Would really appreciate any guidance here. I know I'm just bashing my head against something really simple.
Thanks for your input.
Solutions Both @Desktop-Firework and @Kaushal-Sharma's solutions worked well, in different ways. I also wanted to add the simplest way to do what I was originally trying to do once I got it to work based on their answers. It's obvious to anyone above my level of experience with Python, but maybe it'll help someone in my situation down the line.
I was simply doing an "if" to check if an index point existed, when I obviously should have been using a try-except.
So, using my original code, I could solve what I needed by simply changing:
if url.split('/')[5]:
continue
into
isPath = 1
try: link.split("/")[5]
except IndexError: isPath = 0
Just adding this as it directly answers what I was trying to do at its most basic element. It is not as robust or elegant as either of the provided solutions from the other contributors, obviously.