I have pyspark code that decodes url encoded string:
df.withColumn("clean_url", F.expr("""reflect("java.net.URLDecoder", "decode", url, "UTF-8")"""))
But sometimes it gets string with illegal characters and throws an exception:
Caused by: java.lang.IllegalArgumentException: URLDecoder: Illegal hex characters in escape (%) pattern - For input string: "\x"
How can i skip those exceptions and return NULL if string can not be decoded?
I tried pyspark UDF functions to skip exceptions, but on large set of data stumble on OOM errors and performance decline. So prefer using inbuild sql functions. I don't want to rewrite whole code in scala with UDFs to catch that exception. I am hoping to handle it with spark sql functions.