I have a dataframe with around 50+ columns all in the "long" format. I would like to batch process 40 of them to be converted to "integer" format.
Do I have to keep repeating below?
df = df \
.withColumn('colA', col('colA').cast(IntegerType())) \
.withColumn('colB', col('colB').cast(IntegerType())) \
.withColumn('colC', col('colC').cast(IntegerType())) \
....
Above looks pretty manual to me. I am new to PySpark, so not sure if I can put all columns into a list, and only use cast once (like what I would have done in Python).
Greatly appreciate your help!