AWS Databricks Spark - Data Types

Viewed 18

I'm coming from a world of SQL Server, and I'm new to idea of distributed computing offered by Spark.

I can't find rather important answer around wise usage of data types inside of DataFrames.

It's crucial to use smallest possible data type on RDMS world that give great benefit over storage hence speed ups data manipulations.

Yet, out of all books articles everyone seems to totally ignore this subject. Does it really doesn't matter with Spark? No benefit of using smallint (1 byte) over int (4 bytes)? char Vs varchar?

Would appreciate any reference to articles, benchmarks and documentations that refer to variances of performance considering various data types.

Thank You

0 Answers
Related