How to optimize creation of pyspark dataframe from a list of tuples

Viewed 55
import itertools

l1 = [1,2,3,4,5]
l2 = list(itertools.combinations(l1, 2))
print(l2)

newdf = spark.createDataFrame(l2,['record1', 'record2'])
display(newdf)

This is the code that I tried and it works but takes a long time (for e.g. when l1 is of size 50 million). Is there a better and optimized way to do this using pandas or pyspark ?

0 Answers
Related