I'm using Hive 3.1.0 cluster on HDInsights 4.0.
The orc and parquet with the same data were created using spark, with schema (a string, b int, c string).
create external table a_st_b_int_d_st_orc(a string, b int, d string) stored as orc location
<path_to_spark_created_files>select * from
a_st_b_int_d_st_orc;+----+----+------+ | a | b | d | +----+----+------+ | 1 | 2 | abc | | 2 | 3 | bcd | +----+----+------+create external table a_st_b_int_d_st_parquet(a string, b int, d string) stored as parquet location
<path_to_spark_created_files>select * from
a_st_b_int_d_st_parquet;+----+----+-------+ | a | b | d | +----+----+-------+ | 1 | 2 | NULL | | 2 | 3 | NULL | +----+----+-------+
The default behavior of hive native ORC-Reader is that it maps meta-store column names by position with orc files.
There were JIRAs created to map columns by name and reverted as well.
The behavior wrt parquet can be configured using parquet.column.index.access although default being column resolution by name.
In Presto also we can specify hive.orc.use-column-names=true
How to turn this off the default ORC behavior in hive?