Apache pyspark remove stopwords and calculate

Viewed 210

I have the following .csv file (ID, title, book title, author etc):

csv

I want to compute all the n-combinations (from each title I want all the 4-word combinations) from the titles (column 2) of the articles (with n=4), after I remove the stopwords.

I have created the dataframe:

df_hdfs = sc.read.option('delimiter', ',').option('header', 'true')\.csv("/user/articles.csv")

I have created an rdd with the titles column:

rdd = df_hdfs.rdd.map(lambda x: (x[1]))

and it seems like this:

rdd

Now, I realize that I have to tokenize each string of RDD into words and then remove the stopwords. I would need a little help on how to do this and how to compute the combinations.

Thanks.

0 Answers
Related