I have a text document. I need to find the possible counts of repeating word pairs in the overall document. For example, I have the below word document. The document has two lines, each line separated by ';'. Document:
My name is Sam My name is Sam My name is Sam;
My name is Sam;
I am working on pairwords count.The expected out is:
[(('my', 'my'), 3), (('name', 'is'), 7), (('is', 'name'), 3), (('sam', 'sam'), 3), (('my', 'name'), 7), (('name', 'sam'), 7), (('is', 'my'), 3), (('sam', 'is'), 3), (('my', 'sam'), 7), (('name', 'name'), 3), (('is', 'is'), 3), (('sam', 'my'), 3), (('my', 'is'), 7), (('name', 'my'), 3), (('is', 'sam'), 7), (('sam', 'name'), 3)]
If I use:
wordPairCount = rddData.map(lambda line: line.split()).flatMap(lambda x: [((x[i], x[i + 1]), 1) for i in range(0, len(x) - 1)]).reduceByKey(lambda a,b:a + b)
I get pair-words of consecutive words and their count of re-occurences.
How can I pair each word with every other word in the line and then search for the same pair in all lines?
Can someone please have a look? Thanks