I am trying to read many small avro files(~2.8 MB) that are coming from kafka feed per min. into pyspark dataframe using df_ss7 = spark.read.format('avro').load('gs:<gcs_path>/2022/02/13/*/*/*.avro').
The problem here is when I hit enter it takes very long time to load the as there are huge small files. I know this can be dealt with merging the small file in small number of very large files but I have no idea how that can be done in GCS buckets. I am proving the following parameters to my pyspark shell
--executor-cores 5 --executor-memory 8g --driver-memory 8g
Note - I am using GCP Dataproc cluster to process data using pyspark.