I need to split the words based on the character '/' and reform the words in this way:
This dataframe contains some kids and their presents for Easter. Some kids have two presents, while some have only one.
data = {'Presents':['Pink Doll / Ball', 'Bear/ Ball', 'Barbie', 'Blue Sunglasses/Airplane', 'Orange Kitchen/Car', 'Bear/Doll', 'Purple Game'],
'Kids': ['Chris', 'Jane', 'Betty', 'Harry', 'Claire', 'Sofia', 'Alex']
}
df = pd.DataFrame (data, columns = ['Presents', 'Kids'])
print (df)
This dataframe looks like this:
Presents Kids
0 Pink Doll / Ball Chris
1 Bear/ Ball Jane
2 Barbie Betty
3 Blue Sunglasses/Airplane Harry
4 Orange Kitchen/Car Claire
5 Bear/Doll Sofia
6 Purple Game Alex
I try to delimit their presents and also to reform them in this way, keeping their associated colors:
'Pink Doll/Ball' will be split into two parts: 'Pink Doll', 'Pink Ball'. In addition to this, the same kid should be associated to their presents.
The colours and the presents can be anything, we just know that the structure is: Colour Present1/Present2, or Colour Present or just Present. So finally, it should be:
- for Colour Present/Present --> Colour Present1 and Colour Present2
- for Colour Present ---> Colour Present
- for Present ---> Present
So the final dataframe should look like this:
Presents Kids
0 Pink Doll Chris
1 Pink Ball Chris
2 Bear Jane
3 Ball Jane
4 Barbie Betty
5 Blue Sunglasses Harry
6 Blue Airplane Harry
7 Orange Kitchen Claire
8 Orange Car Claire
9 Bear Sofia
10 Doll Sofia
11 Purple Game Alex
My first approach was to transform the columns into lists and work with lists. Like this:
def count_total_words(string):
total = 1
for i in range(len(string)):
if (string[i] == ' '):
total = total + 1
return total
coloured_presents_to_remove_list = []
index_with_slash_list = []
first_present = ''
second_present= ''
index_with_slash = -1
refactored_second_present = ''
for coloured_present in coloured_presents_list:
if (coloured_present.find('/') >= 0):
index_with_slash = coloured_presents_list.index(coloured_present)
index_with_slash_list.append(index_with_slash)
first_present, second_present = coloured_present.split('/')
coloured_presents_to_remove_list.append(coloured_present)
if count_total_words(first_present) == 2:
refactored_second_present = first_present.split(' ', 1)[0] + ' ' + second_present
second_present = refactored_second_present
coloured_presents_list.append(first_present)
coloured_presents_list.append(second_present)
kids_list.insert(coloured_presents_list.index(first_present), kids_list[index_with_slash])
kids_list.insert(coloured_presents_list.index(second_present), kids_list[index_with_slash])
for present in coloured_presents_to_remove_list:
coloured_presents_list.remove(present)
for index in index_with_slash_list:
kids_list.pop(index)
However, I have realized that in some point, I might lose some index by mistake so I tried working with pandas into dataframe.
mask = df['Presents'].str.contains('/', na=False, regex=False)
df['First Present'], df['Second Present'] = df.loc[mask, 'Presents'].split('/')