How to register UDF (python and scala) when using 'spark-sql' submit?

Viewed 232

It's possible to register UDF with code, in the context, before using sql api. Spark proposes a command line tool: spark-sql to submit some SQL requests.

This tool use spark-submit with --class org.apache.spark.sql.hive.thriftserver.SparkSQLCLIDriver.

It's not possible to register a UDF before using spark-sql, but it's possible to add some jar or py-files.

What are the ways to use spark-sql with some registered functions?

1 Answers

In general, JVM (Java / Scala) classes implementing the following interfaces:

  • org.apache.hadoop.hive.ql.exec.{UDF, UDAF}
  • org.apache.hadoop.hive.ql.udf.generic.{AbstractGenericUDAFResolver, GenericUDF, GenericUDTF}
  • org.apache.spark.sql.expressions.UserDefinedAggregateFunction

can be registered using CREATE FUNCTION interface. Permanent functions will be available for all sessions.

There is no such capability for other UDF variants, including Python UDFs, ATM.

Related