import itertools
l1 = [1,2,3,4,5]
l2 = list(itertools.combinations(l1, 2))
print(l2)
newdf = spark.createDataFrame(l2,['record1', 'record2'])
display(newdf)
This is the code that I tried and it works but takes a long time (for e.g. when l1 is of size 50 million). Is there a better and optimized way to do this using pandas or pyspark ?