spark.table vs sql() AccessControlException

Viewed 966

Trying to run

  spark.table("db.table")
    .groupBy($"date")
    .agg(sum($"total"))

returns

org.apache.spark.sql.AnalysisException: org.apache.hadoop.hive.ql.metadata.HiveException: Unable to alter table. java.security.AccessControlException: Permission denied: user=user, access=WRITE, inode="/sources/db/table":tech_user:bgd_group:drwxr-x---

the same script but written as

  sql("SELECT sum(total) FROM db.table group by date").show()

returns actual result.

I don't understand why this is happening. What is the first script trying to write exactly? Some staging result? I have read permission for this table and I'm only trying to perform some aggregations. Using Spark 2.2 for this.

1 Answers

In Spark 2.2, the default for spark.sql.hive.caseSensitiveInferenceMode was changed from NEVER_INFER to INFER_AND_SAVE. This mode causes Spark to infer (from underlying files) and try to save case-sensitive schema into Hive metastore. This will fail if the user executing the command wasn't granted permissions to update HMS.

Obvious workaround is to set inference mode back to NEVER_INFER, or INFER_ONLY if application relies on column names as they present in files (CaseSensitivE).

Related