I have the following df:
Score num_comments titles
0 134 518 Uhaul implement nicotine-free hiring policy
1 28 43 Orangutan saves child from a giant volcano
2 30 114 Swimmer dies in a horrific shark attack in harbour
3 745 298 More teenagers than ever are addicted to glue
4 40 67 Lebanese lawyers union accuse Al Capone of fraud
...
9366 345 32 City of Louisville closed off this summer
9367 1200 234 New york rats "stronger than ever", reports say
9368 432 123 Congolese militia shipwrecked in Norway
9369 594 203 Scientists now agree on how to use ice in drinks
9370 611 153 Historic drought hits Atlantis
Now I would like to create a new dataframe where I can see what score and how many comments each word gets. Like this: df2
Word score num_comments
Uhaul 134 518
implement 134 518
nicotine-free 134 518
hiring 134 518
policy 134 518
Orangutan 28 43
saves 28 43
child 28 43
from 28 43
a 28 43
giant 28 43
volcano 28 43
...
etc..
I have tried Splitting the title into separate words and then exploding:
In [9]: df3
Out[9]:
df3['titles_split'] = df['titles'].str.split()
This gave me a column that looked like this:
Score num_comments titles_split
0 134 518 [Uhaul, implement, nicotine-free, hiring, policy]
1 28 43 [Orangutan, saves, child, from, a, giant, volcano]
2 30 114 [Swimmer, dies, in, a, horrific, shark, attack, in, harbour]
3 745 298 [More, teenagers, than, ever, are, addicted, to, glue]
4 40 67 [Lebanese, lawyers, union, accuse, Al, Capone, of, fraud]
...
9366 345 32 [City, of, Louisville, closed, off, this, summer]
9367 1200 234 [New, york, rats, stronger, than, ever, reports, say]
9368 432 123 [Congolese, militia, shipwrecked, in, Norway]
9369 594 203 [Scientists, now, agree, on, how, to, use, ice, in, drinks]
9370 611 153 [Historic, drought, hits, Atlantis]
Then I tried this code:
df3.explode(df3.assign(titles_split=df3.titles_split.str.split(',')), 'titles_split')
But I got the following error message:
ValueError: column must be a scalar, tuple, or list thereof
The same thing happened when I tried it for titles in df2.
I also tried creating new columns that repeated scores and num_comments as many times as there are words in titles (or titles_split). The idea was to create a dataframe like this:
In [9]: df4
Out[9]:
Score num_comments titles_split score_repeated
0 134 518 [Uhaul, implement, nicotine-free, hiring, policy] 134,134,134,134,134,134
1 28 43 [Orangutan, saves, child, from, a, giant, volcano] 28,28,28,28,28,28,28
2 30 114 [Swimmer, dies, in, a, horrific, shark, attack, in, harbour] 30,30,30 etc..
3 745 298 [More, teenagers, than, ever, are, addicted, to, glue] etc.
4 40 67 [Lebanese, lawyers, union, accuse, Al, Capone, of, fraud] etc
...
9366 345 32 [City, of, Louisville, closed, off, this, summer] etc
9367 1200 234 [New, york, rats, stronger, than, ever, reports, say] etc
9368 432 123 [Congolese, militia, shipwrecked, in, Norway] etc
9369 594 203 [Scientists, now, agree, on, how, to, use, ice, in, drinks] etc
9370 611 153 [Historic, drought, hits, Atlantis] etc
And then explode on titles_split, score_repeated and comments_repeated like this:
df4.explode(['titles_split', 'score_repeated', 'comments_repeated'])
But I never got to that point because I couldn't get repeated columns. I tried the following code:
df3['score_repeat'] = df3.apply(lambda x: [x.score] * len(x.titles_split) , axis =1)
Which gave me this error message:
TypeError: object of type 'float' has no len()
Then I tried:
df3['score_repeat'] = [[y] * x for x, y in zip(df3['titles_split'].str.len(),df['score'])]
Which gave me:
TypeError: can't multiply sequence by non-int of type 'float'
But I am not even sure I am going about this the right way. Do I even need to create score_repeated and comments_repeated?