remove the string which has lower number of string

Viewed 99
 column   1      2           3        4
     A     hot   too hot    school   playground
     B    weather cold     too cold  jacket
     C    rain    water    dirt     rain coat
   

As you can see, there are repeated strings in different columns. Like hot and too hot. Is there a way where I can keep the string with longer length and delete that specific cell which has the same string?

the output I would want is something like this:

 Column   1         2           3          4
    A             too hot    school    playground
    B    weather             too cold   jacket
    B              water        dirt     rain coat

data['repeat'] = data[['1', '2','3','4']].apply(lambda x: x.str.contains('hot'))

This is the code I am working on but this too is only allows me to select a specific string, which is not good if I am working on a large dataset.

1 Answers

My two cents of a solution.

Loop the columns; split the string in the columns into words and loop over the words and check if the words occur in the strings in the other columns; if the case add the index of the column of the shorter string to a set of ids where the string needs to be emptied. Finally empty the string in the columns if the index is in the set.

def remove_strings(col_string):
    delete_ids = set()
    for i, s1 in enumerate(col_string):
        words = s1.split()
        for word in words:
            for j, s2 in enumerate(col_string):
                if i != j and word in s2:
                    if len(s1) >= len(s2):
                        delete_ids.add(j)
                    else:
                        delete_ids.add(i)

    for id in delete_ids:
        col_string[id] = ''

    return col_string

col_strings = []
col_strings.append(['hot', 'too hot', 'school', 'playground'])
col_strings.append(['hot chocolate', 'too hot', 'school', 'playground'])
col_strings.append(['weather', 'cold', 'too cold', 'jacket'])
col_strings.append(['rain', 'water', 'dirt', 'rain coat'])

for col_string in col_strings:
    print(f'original: {col_string}')
    print(f'string removed: {remove_strings(col_string)}')

original: ['hot', 'too hot', 'school', 'playground']
string removed: ['', 'too hot', 'school', 'playground']
original: ['hot chocolate', 'too hot', 'school', 'playground']
string removed: ['hot chocolate', '', 'school', 'playground']
original: ['weather', 'cold', 'too cold', 'jacket']
string removed: ['weather', '', 'too cold', 'jacket']
original: ['rain', 'water', 'dirt', 'rain coat']
string removed: ['', 'water', 'dirt', 'rain coat']
Related