HDFS vs HIVE partitioning

Viewed 4431

This may be a simple thing but i'm struggling to find the answer. When the data is loaded to HDFS its distributed and loaded into multiple nodes. The data is partitioned and distributed.
For HIVE there is a separate option to PARTITION the data. I'm pretty sure that even if you don't mention the PARTITION option, the data will be split and distributed to different nodes on the cluster, when loading a hive table. What additional benefit does this command give in this case.

2 Answers

summarizing comments and for Hadoop v1-v2.x:

  • a logical partitioning, eg. related to a date or field in a string, as written in the comments above, is only possible in hive, hcat or a another sql or parallel engine working on top of hadoop, using a fileformat which supports partitioning (Parquet, ORC, CSV are ok, but eg. XML is hard or nearly impossible to partition)

  • logical partitioning (like in hive, hcat) can be used as a replacement for not having an indexes

  • 'partitioning of hdfs storage' on local or distributed nodes is possible by defining the partitions during setup of hdfs, see https://docs.hortonworks.com/HDPDocuments/HDP2/HDP-2.6.5/bk_cluster-planning/content/ch_partitioning_chapter.html

  • HDFS is able to "balance" or 'distribute' blocks over nodes

  • Natively, blocks can't be split and distributed to folders by HDFS according to their content, only moved at whole to another node

  • blocks (not files!) are replicated in the HDFS cluster according to the HDFS replication factor:

    $ hdfs fsck /
    

(thanks David and Kris for your discussion above, also explains most of it and please take this post as summary)

Related