I'm trying to create a new sample based on some other samples but I'm doing/understanding something wrong. I have 34 samples which I assume is relatively lognorm distributed. Based on this samples I want to generate 2000 new samples. Here is the code that I'm running:
import numpy as np
from scipy import stats
import matplotlib.pyplot as plt
samples = [480, 900, 1140, 1260, 1260, 1440, 1800, 1860, 1980, 2220, 2640, 2700,
2880, 3420, 3480, 3600, 3840, 4020, 4200, 4320, 4380, 4920, 5160,
5280, 6900, 7680, 9000, 10320, 10500, 10800, 15000, 21600, 25200,
39000]
plt.plot(samples, 1 - np.linspace(0, 1, len(samples)))
std, loc, scale = stats.lognorm.fit(samples)
new_samples = stats.lognorm(std, loc=loc, scale=scale).rvs(size=2000)
a = plt.hist(new_samples, bins=range(100, 40000, 200),
weights=np.ones(len(new_samples)) / len(new_samples))
plt.show()
Here is the plot, and as you can see there are really few samples above 1000, although the sample contained rather many above 1000.
How do I best generate a sample that better represent the expected values?




