Schema resolution issue in Impala when you add a new column in table at start or in between

Viewed 848

The issue i'm facing is related to schema resolution in Impala. Impala currently does not support resolving schema-to-file metadata by name - it does so only by index i.e Impala looks up columns within a Parquet file based on the order of columns in the table.

For example:

Table T1, with 2 columns [A: int and B:String]

A|B
1|foo
2|bar

Added new column [X: string]

  X | A | B
test| 3 |bar1

Impala read this table based on the index. So at index(0), there are 2 different data type -

  • string(w.r.t column X in NEW data),
  • int(w.r.t column A in OLD data).

In my scenario, I had Re-created the table using the updated DDL to accommodate the new column, but since I added the new column in the middle...I'm getting a error when I query it.

File 'hdfs://nameservice1/user/hive/warehouse/ra/rds_data.db/slot_data/date=20210107/part-00002.c000' has an incompatible Parquet schema for column 'rds_data.slot_data.slotnumber'. Column type: INT, Parquet schema: optional byte_array location

The workaround that I found were:

1. Use a schema resolution property

Set PARQUET_FALLBACK_SCHEMA_RESOLUTION=name

The PARQUET_FALLBACK_SCHEMA_RESOLUTION query option allows Impala to lookup columns within Parquet files by column name, rather than column order. But this is required to be hit in every HUE/impala session by every user(in my org.) who will query my table.

2. Overwriting data

Since the the data in the new created table is being populated constantly via spark jobs...I'll have to move new column X at the end of the table and then overwrite the data that came after I executed the updated DDL(last 2-3 days).

Currently solution #2 is the only permanent solution that I can think...Is there any other solution/workaround in which we don't have to touch/overwrite the data and still it works for all the users in my Org?

0 Answers
Related