Spark UDF now throws ArrayIndexOutOfBoundsException

Viewed 135

I wrote a UDF in Spark (3.0.0) to do an MD5 hash of columns that looks like this:

def Md5Hash(text: String): String = {
  java.security.MessageDigest.getInstance("MD5")
    .digest(text.getBytes())
    .map(0xFF & _)
    .map("%02x".format(_))
    .foldLeft("") { _ + _ }
}
val md5Hash: UserDefinedFunction = udf(Md5Hash(_))

This function has worked fine for me for months, but it is now failing at runtime:

org.apache.spark.SparkException: Failed to execute user defined function(UDFs$$$Lambda$3876/1265187815: (string) => string)
....
Caused by: java.lang.ArrayIndexOutOfBoundsException
        at sun.security.provider.DigestBase.engineUpdate(DigestBase.java:116)
        at sun.security.provider.MD5.implDigest(MD5.java:109)
        at sun.security.provider.DigestBase.engineDigest(DigestBase.java:207)
        at sun.security.provider.DigestBase.engineDigest(DigestBase.java:186)
        at java.security.MessageDigest$Delegate.engineDigest(MessageDigest.java:592)
        at java.security.MessageDigest.digest(MessageDigest.java:365)
        at java.security.MessageDigest.digest(MessageDigest.java:411)

It still works on some small datasets, but I have another larger dataset (10Ms of rows, so not terribly huge) that fails here. I couldn't find any indication that the data I'm trying to hash are bizarre in any way -- all input values are non-null, ASCII strings. What might cause this error when it previously worked fine? I'm running in AWS EMR 6.1.0 with Spark 3.0.0.

0 Answers
Related