First, please allow me to start by saying that I am pretty new to Spark-SQL.
I am trying to understand various Join types and strategies in Spark-Sql, I wish to be able to know about an approach to approximate the sizes of the tables (which are participating in a Join, aggregations etc) in order to estimate/tune the expected execution time by understanding what is really happening under the hood to help me to pick the Join strategy which is best suited for that scenario (in Spark-SQL through hints etc).
Of course, the table row-counts offers a good starting point, but I want to be able to estimate the sizes in terms of bytes/KB/MB/GB/TBs, to be cognizant which table would/would not fit in memory etc) which in turn would allow me to write more efficient SQL queries by choosing the Join type/strategy etc that is best suited for that scenario.
Note : I do not have access to PySpark. We query the Glue tables thru Sql Workbench connected to the Glue catalog with a Hive jdbc driver.
Any help is appreciated.
Thanks.