I have a external source file news.txt of type:
20030249,the old men
20040229,I like the way school and teachers work
20050249,another title goes here for any reason
20060269,text and strings are similar
20070551,cowbows love to ride horses
and a list words.txt of words:
the
a
school
horses
The following code creates pairs RDDs in the form of
['2003', ['the', 'old', 'men'], ['2004', ['I', 'like', 'the',...
After the following code, I would like to add an RDD pair transformation code to remove the words words.txt from the values list of the RDD "pair":
source = sc.textFile("news.txt")
stopwords = sc.textFile("words.txt")
pair = source.map(lambda s: [s[0:4],s[9::].split(' ')])
I have tried several in vain but I'm sure I'm close:
pair1 = pair.filter(lambda x: x not in stopwords)
pair1 = pair.map(lambda ws: for w in ws if w not in stopwords)
pair1 = pair.filter(lambda a: a != stopwords)
pair1 = pair.mapValues(lambda x: x not in stopwords)