How to create a SPACE between word and open bracket in list of sentences

Viewed 231

In below list, there are actually two dupes. But due to difference of SPACE between second word of sentence and (, its treating them as unique sentences.

By using Python - Regular Expressions, how to create addition space between words. (example: 1st item) 'United States(US)', should be changed to 'United States (US)' (same as 2nd item)

listx = 
['United States(US)',
 'United States (US)',
 'New York(NY)',
 'New York (NY)']

Expected Output list is

['United States (US)',
 'United States (US)',
 'New York (NY)',
 'New York (NY)']

Actually, i am trying eliminate duplicate sentences from the list and considering this is one of the approach by making the sentences similar first. Please suggest.

3 Answers

You can search for a letter immediately followed by an open parentheses

>>> [re.sub(r'(\w)\(', r'\1 (', i) for i in listx]
['United States (US)',
 'United States (US)',
 'New York (NY)',
 'New York (NY)']

To remove the duplicates you can create a set from this generator expression

>>> set(re.sub(r'(\w)\(', r'\1 (', i) for i in listx)
{'United States (US)', 'New York (NY)'}

You can try this. You can use re.sub here.

listx = ['United States(US)', 'United States (US)', 'New York(NY)', 'New York (NY)']

[re.sub(r'.(\(.*\))',r' \1',i) for i in listx]
# ['United State (US)', 'United States (US)', 'New Yor (NY)', 'New York (NY)']

Regex pattern explanation:

  • . to match any character
  • ( start of group bracket
  • \( match (
  • .* match greedily.
  • ' \1' sub the match group with space the matched group.
  • regex live demo

You can do

    new_listx = ["{} {}".format(re.match('(.*)(\(.*\))', i).group(1).rstrip() ,re.match('(.*)(\(.*\))', i).group(2)) for i in listx]
    print(new_listx)

Output

['United States (US)', 'United States (US)', 'New York (NY)', 'New York (NY)']

The regex is splitting the text to 2 groups one is before the () and the second in the () after that it's trimming the space form the right of the first group.
then you can do

print(set(new_listx))

and you will get a unique values set.

{'New York (NY)', 'United States (US)'}
Related