How to find the best fitting parametric distribution for an empirical dataset (stock returns)?

Viewed 178

Given some real-valued empirical data (time series), I could convert it to a histogram to have an (non-parametric) empirical distribution of the data, but histograms are blocky and jagged.

Instead, I would like to identify the best-fitting parametric distribution from the scipy or scipy.stats libraries of distribution functions, so that I can artificially generate a parametric distribution that closely fits the empirical distribution of my real data.

If the empirical data are monthly returns of empirical AAPL stock returns, for example, I know that the parametric Johnson-SU distribution resembles, and can mimic, stock return distributions because of its customizable skew. However, the Johnson SU distribution in scipy requires four input parameters to be calibrated. How can I search for the best parameter settings of this parametric distribution from scipy that fits to the empirical distribution of my sample of AAPL returns?

1 Answers

Q : "I would like to identify the best-fitting parametric distribution from the scipy or scipy.stats libraries of distribution functions, so that I can artificially generate a parametric distribution that closely fits the empirical distribution of my real data."

The link from @SeverinPappadeux above might help ( K-S tests are fine ) yet it serves well but for the analytical comparison of a pair of already complete distribution, not for the process of the actual constructive generation thereof.

So let's disambiguate the goal:
- is the task focused on using scipy / scipy.stats generators?
or
- is the task focused on achieving a process of generating a synthetic distributions well-enough matching the empirical "original"?


Should the former is your wish,
then
we run into an oxymoron, to seek a parametrise-able (scripted) distribution generator-engine, that will (in some sense of a "best"-ness) match a principally un-scriptable empirical distribution
well, as one might still wish to do so
then
you will indeed end up in some sort of a painful ParameterSPACE search strategy (using the ready-made or customised scipy/scipy.stats hardcoded-generators) that will try to find the "best"-matching values of the ParameterSPACE-vector of these generators' hard-coded parameters. This may teach you to some degree about the sin of growing dimensionality ( the more parameters a hard-coded generator has, the larger is the ParameterSPACE search-space, going into O( n * i^N * f^M * c^P * b^Q) double trouble, having N-integer, M-float, P-cardinal and Q-boolean parameters of a respective hard-coded generator, which goes pretty nasty against your time-budget, doesn't it? ).


Should the latter is the case,
then
we may focus on a more productive way by proper defining what is the "wellness"-of-"matching" the "original".

The first candidate for this is to generate a pretty random ( quite easily a PRNG-produced ) noise, that if not too "strong" inside the PriceDOMAIN direction may be simply added to the empirical-"original" & here we go.

More sophistication might be added, using the same trick of using superposition, drop-out(s), frequency-specific tricks, outlier add-on(s) (if later testing properties/limits of robustness of some dataflow-responsive strategies et al)

Anyway, all these methods for the latter target have a lovely property of not going anywhere wild into any vast searches of high-dimensionality ParameterSPACEs, but are often as nice as just O( n )-scaled -- that's cool, isn't it?

So, just one's own imagination is the limit here :o)

Related