I am trying to create a simulated dataset of emails. For that, I want to generate recipients based on 2 parameters:
- how many recipients
- how many domains those recipients will be from
For that, I have created a dataframe whose first few rows are as follows:
import pandas as pd
data = {'Date':['19/06/2022', '19/06/2022', '20/06/2022', '20/06/2022', '21/06/2022', '21/06/2022'],
'Time':['8:25:12', '9:21:33', '18:26:28', '11:39:23', '7:30:47', '20:27:48'],
'Sender': ['fqqp@abc.com', 'jald@abc.com', 'acpo@abc.com', 'smfa@abc.com', 'jald@abc.com', 'fczz@abc.com'],
'Number of recipient domains': [2, 3, 3, 4, 3, 5],
'Number of recipients': [7, 4, 7, 4, 6, 7]
}
df = pd.DataFrame(data)
Now I need to generate the Recipients column, that will hold some random recipients according to the columns Number of recipient domains and Number of recipients. For example, take the first row - it should generate 7 email addresses from 2 domains (like, @abc.com and @xyz.pqr for instance).
How do I do that?
I can write a function to generate email addresses with only the number of email addresses to be generated as argument (i.e., with the number of domains parameter removed):
import string, random
def random_email_gen(num_recipients):
emails = []
for _ in range(int(num_recipients)):
name = ''.join(random.choice(string.ascii_lowercase) for _ in range(4))
domain = ''.join(random.choice(string.ascii_lowercase) for _ in range(3))
emails.append(name + '@' + domain + '.com')
return emails
random_email_gen(num_recipients=7)
>>> ['dvxh@jnz.com',
'anpd@tvl.com',
'nons@voz.com',
'fneu@vcg.com',
'tqng@nnm.com',
'xlib@lzv.com',
'copy@jff.com']
But how do I extend it to generate the randomized email addresses with the number of domains parameter as well?
