Here is a sampling method. I tried:
sample=2000
sample_df = df.groupby('prefix').sample(n=sample, random_state=1)
It groups df by prefix and for each group, it samples 2k items. I have 9 groups. I want to sample 18k but weighted by the number in each group.