Given an RDD with tuples like
(key1, 0)
(key2, 5)
(key3, 11)
(key1, 44)
(key2, 0)
(key3, 43)
(key1, 0)
(key2, 5)
(key3, 33)
For each key, I wish to count 2 values, total count of values per key, or regular output of countByKey(), and second count, count of positive numbers by key.
So the result would be like:
[(key1, 3, 1),
(key2, 3, 2),
(key3, 3, 3)]
I would like to use only map and reduceByKey functions:
def map(value):
return (value[0], (value[1], 1, 1))
return a key value pair, where value is the triple of the numerical value, and two integers used for counting
def reduce(val1, val2):
# if value[0] is positive, increment first counter
# in any case, always increment second counter
data.map(map).reduceByKey(reduce).collect()
would then be something like:
[(key1, 3, 1),
(key2, 3, 2),
(key3, 3, 3)]