how to compare hadoop result

Viewed 36

I am writing a map reduce program to find the file that contains the most words.

Now, I am able to use map reduce to find the number of words contained in each file. However, I am unsure how I can store the number of words in each file and then compare it and find the file that contains the most words using map reduce.

My idea so far:

Having several jobs to find the number of words in each file, like this

file_name | number of words 
file_1      5
file_2      10
file_3      15

then start another job, in the reducer find the maximum number of words and finally get the following result

file_3

I wonder: does the approach make sense? Is there any other way to find the file that contains the most words via map reduce?

1 Answers

Finding minimum/maximum in mapreduce isn't a good use case for it. You need to force data to one reducer

For example, iterate the input, and write from the mapper

null, file_1=5
null, file_2=10
null, file_3=15

Then, iterate the values in the reducer and find the maximum value like you would for any array. You'd need to split out the delimiter and you have both the filename and the "number of words"

Related