I am coding Map Reduce paradigm and it is working fine in command prompt as .py
#map reduce paradigm
#mapper
data=["This is delhi\n", "This is paris\n", "This is Moscow\n"]
#create a new file and leave it open to write over it
file = open("jose.txt","w+")
for line in data:
#remove leading and trailing whitespace - choosing every line
line=line.strip()
#split the line into words
words=line.split()
#increase counters
for word in words:
file.write("%s\t%s" % (word,1))
file.write("\n")
print("%s\t%s" % (word,1))
file.close()
#Map reduce paradigm
#reducer
from operator import itemgetter import sys
current_word = None current_count = 0 word = None
file = open("jose.txt","r") lines = file.readlines() lines.sort()
# input comes from STDIN for line in lines:
# remove leading and trailing whitespace
line = line.strip()
# parse the input we got from mapper.py
word, count = line.split()
# convert count (currently a string) to int
count = int(count)
if current_word == word:
current_count += count
else:
if current_word:
# write result to STDOUT
print ('%s\t%s' % (current_word, current_count))
current_count = count
current_word = word
# do not forget to output the last word if needed!
if current_word == word:
print ('%s\t%s' % (current_word, current_count))
file.close()
The output is
This 1
is 1
delhi 1
This 1
is 1
paris 1
This 1
is 1
Moscow 1
Moscow 1
This 3
delhi 1
is 3
paris 1
It is good, but my professor is asking now:
Generate .py and execute it from command prompt, typing only "python mapper.py". It is easy, I did it.
Generate two .py files, one mapper the other one reducer, and execute them typing only "python reducer.py mapper.py" I cannot do it.
In some way the output of mapper.py have to be the input of reducer.py. Basically is the same code than in step 1, but I have to split it and then link the output of mapper.py with the input of reducer.py