How to filter out lines starting with 'URL' in filter pyspark RDD

Viewed 145

I have a pyspark sc initialized.

task1 = (text.filter(lambda x: len(x)>0 )) # to filter empty lines
task1.collect()

My goal is to filter out lines starting with 'URL' in this text snippet:

['URL: http://www.nytimes.com/2016/06/30/sports/baseball/washington-nationals-max-scherzer-baffles-mets-completing-a-sweep.html', 'WASHINGTON — Stellar pitching kept the Mets afloat in the first half of last season despite their offensive woes.

How can I do this in pyspark syntax easily?

3 Answers

you can use regex

import re

reg = re.compile('^(?!URL).*')
task1 = text.filter(lambda x: reg.match(x))

Question needed sample input and output. I assume the data provided are rows in a table. If that's not the case, happy to change answer after clarification. If it is then;

Say data is;

+---+--------------------+
|SID|           Attribute|
+---+--------------------+
|  1|[URL: http://www....|
|  2|scherzer-baffles-...|
|  3|kept the Mets afl...|
+---+--------------------+

let us use filter alongside PySpark expr(); a SQL function to execute SQL-like expressions in data frames

from pyspark.sql.functions import *
df.filter(expr("Attribute like '[__%'")).show()#Finds any values that start with "[" and are at least 3 characters in length

+---+--------------------+
|SID|           Attribute|
+---+--------------------+
|  1|[URL: http://www....|
+---+--------------------+

If you're already splitting the file into lines (which is likely), you could probably use:

task2 = text.filter(lambda x: x[0:3] != 'URL')
Related