I have a Spark application that writes output files in Avro format. Now I would like that data to be available in Hive, because an application which would utilise that data can only do so through a Hive table.
It is described here that one can do that by using CREATE EXTERNAL TABLE in Hive. Now my question is, how efficient is the CREATE EXTERNAL TABLE method. Would it copy all the Avro data somewhere else on the HDFS to work, or does it just create some metainfo, which it can use to query Avro data?
Also, what if I want to keep on adding new Avro data to that table. Can I create such an external table once, and then keep adding the new Avro data to it? Also what if someone queries the data while it's being updated. Does it allow atomic transactions?