I have a Pyspark DataFrame, I want to randomly sample (From anywhere in the entire df) ~100k unique ID's. The DF is transaction based so an ID will appear multiple times, I want to get 100k distinct ID's and then get all the transaction records for each of those ID's from the DF.
I've tried:
sample = df.sample(False, 0.5, 42)
sample = sample.distinct()
Then I'm unsure how to match It back to the original Df, Also some of the ID's are not clean, I want to be able to put some condition in the sample that says the ID must be for example 10 digits.