How to figure out rising words (or words which are becoming more popular from a list)

Viewed 65

I was wondering how I can figure out rising words (sort of similar to reddit's rising threads sort option). What I mean by rising is, what is becoming popular. What's rising to the top the fastest. Example:

PS: the -, ^ and v are just how the words moved. Also the (a) and (b) are not part of the word. I am just showing their uniqueness this way

At 10:00am I have this list, in this rank

1. Cool    (a)  -
2. Best         -
3. Cool    (b)  -
4. Radical (a)  -
5. Sweet   (a)  -
6. Sweet   (b)  -
7. Radical (b)  -

Then at 10:15am (15 minutes later), the order of the list changed.

1. Best         ^
2. Cool    (a)  v
3. Radical (a)  ^
4. Sweet   (a)  ^
5. Cool    (b)  v
6. Radical (b)  ^
7. Sweet   (b)  v

Then at 10:30am (15 minutes later), the order of the list changes again.

1. Best         -
2. Radical (a)  ^
3. Sweet   (a)  ^
4. Cool    (a)  v
5. Radical (b)  ^
6. Sweet   (b)  ^
7. Cool    (b)  v

As you can see, the word Cool as a whole is clearly the dropping in popularity. Currently my algorithm (I feel, is fairly stupid, but I can't think of any other way).

The way I am doing it now is:

  1. For every word on the list, I count how many ranks it moved up (+ num) or down (- num) or if it didn't move 0. This technically gives me a rate. Ranks moved per 15 minutes
  2. Then if that same word exists twice (like the word Cool), then I average the rate.
  3. Then I sort it from highest to lowest and I have my rising words.

Though I feel this isn't very good (or even makes any sense). It surely doesn't take into account any historical data either, only the new data it receives every 15 minutes.

My question is, how can I figure out the top rising word, bottom rising word and all the words in between.

1 Answers

If you want to take into account historical data you can use some function to decrease weight of data for older changes. That function could be exponent, it decays very fast:

risingRate = 0
for i = 0:n
    risingRate += e^(-i) * RankChange(curr - i)
end
return risingRate

This code uses n + 1 last records for a word to compute its rising rate.

Coefficients on every step would be:

 0: 1
 1: 0.367879
 2: 0.135335
 3: 0.0497871
 4: 0.0183156
 5: 0.00673795
 6: 0.00247875
 7: 0.000911882
 8: 0.000335463
 9: 0.00012341
10: 4.53999e-05

These coefficients assign bigger weight for the most recent coefficients. Which is what you might want.

You can adjust decay rate by using e^(-alpha*i).

Related