Python matching strings within substrings

Viewed 46

I'm writing a program to take a json formatted file and create a proxy PAC file. One of the challenges I've encountered is that the json file contains a mixture of data which is not neatly organized. I would like to summarize the data like so:

Input data:

www.example.com
*.example.com
example.com
myserver.example.com
server.*.example2.com
server.mydomain1.example2.com
server.mydomain2.example2.com
server.mydomain3.example2.com
example2.com

Output data:

*.example.com
example.com
server.*.example2.com
example2.com

I'm trying to find the most python way to summarize the data. Any ideas? I thought of using regular expressions to help with pattern matching but I imagine they can get complicated quite quickly?

1 Answers

I could only come up with a pretty messy way to do this, but I'll try to explain with comments.

import re
l = ["www.example.com",
     "*.example.com",
     "example.com",
     "myserver.example.com",
     "server.*.example2.com",
     "server.mydomain1.example2.com",
     "server.mydomain2.example2.com",
     "server.mydomain3.example2.com",
     "example2.com"]
# Something can only summarize if it contains a wildcard. Otherwise it won't represent the other elements in the list
summarizable = [domain for domain in l if "*" in domain] 
[url for url in l 
    if not bool( # check to see if url is not represented by any of the wildcards
        [1 for summary in summarizable # escape the ., replace * with re wildcard (.*)
            if bool(re.match(summary.replace('.','\.').replace('*','.*'), url)) ])] + summarizable

returns

['example.com', 'example2.com', '*.example.com', 'server.*.example2.com']

The caveat with this solution: if you have two wildcard urls that can be summarized by each other, they both will appear in the final output.

Related