Experiencing strange error message from PyMC3 - Python 3.8.5

Viewed 396

I recently changed my IDE from Spyder to PyCharm, as I felt Spyder used through Anaconda was somewhat bloated. The code below runs more or less fine in Spyder, but I'm having issues in PyCharm. The code runs smoothly up until the data frame gets put in the model(1.1).Observing the metadata(?) related to the file, it should be suitable for analysis.

Every other places I have looked have been dealing with wrong formatting for file import. Far as I can see, this isn't the case with me, as the data fram works fine prior to the model.

Things I have done with hopes of resolving the isse:

  • Decreased the sample size from 100K to 50K to 10K to 1K
  • Changed the project directory with the actual
  • I reinstalled the following dependencies via conda contra pip: numpy, PyMC3, Theano
  • I increased the VM to 100GB
  • Installed new GPU driver for processing
  • I added preprocessing thinking it might have something to do with that

The error that I keep getting when running the model is:

Traceback (most recent call last):
  File "<string>", line 1, in <module>
  File "C:\Users\Greencom\miniconda3\lib\multiprocessing\spawn.py", line 116, in spawn_main
    exitcode = _main(fd, parent_sentinel)
  File "C:\Users\Greencom\miniconda3\lib\multiprocessing\spawn.py", line 125, in _main
    prepare(preparation_data)
  File "C:\Users\Greencom\miniconda3\lib\multiprocessing\spawn.py", line 236, in prepare
    _fixup_main_from_path(data['init_main_from_path'])
  File "C:\Users\Greencom\miniconda3\lib\multiprocessing\spawn.py", line 287, in _fixup_main_from_path
    main_content = runpy.run_path(main_path,
  File "C:\Users\Greencom\miniconda3\lib\runpy.py", line 264, in run_path
    code, fname = _get_code_from_file(run_name, path_name)
  File "C:\Users\Greencom\miniconda3\lib\runpy.py", line 234, in _get_code_from_file
    with io.open_code(decoded_path) as f:
OSError: [Errno 22] Invalid argument: 'C:\\Users\\Greencom\\OneDrive\\Dokumenter\\Trading\\Quant\\Python Analysis\\<input>'

I cannot see where the file directory has any relevancy to the actual model, so I don't understand the counter argument.

How do I solve this?

from scipy import stats
import arviz as az
import numpy as np
import matplotlib.pyplot as plt
import pymc3 as pm
import seaborn as sns
import pandas as pd
from theano import shared
from sklearn import preprocessing

print ( 'Running on PyMC3 v{}'.format ( pm.__version__ ) )

# data import #
EU50p1d = pd.read_excel (
    "C:\\Users\\Greencom/OneDrive\\Dokumenter\\Trading\\Quant\\Python Analysis\\CURRENCYCOM_EU501Dp.xlsx" )

# checking for 0-values #
EU50p1d.isnull ().sum () / len ( EU50p1d )

# Determining upper & lower value for model
lower = int ( EU50p1d.min () ) if int ( EU50p1d.min () ) == float ( EU50p1d.min () ) else float ( EU50p1d.min () )
upper = int ( EU50p1d.max () ) if int ( EU50p1d.max () ) == float ( EU50p1d.max () ) else float ( EU50p1d.max () )

# Gaussian Inferences #
az.plot_kde ( EU50p1d.close, rug=True )
plt.yticks ( [0], alpha=0 )


### 1.1 Model ##
with pm.Model() as model_g:
    μ = pm.Uniform('μ' , lower=2375.7 , upper=3858.3)
    σ = pm.HalfNormal('σ' , sd=100)
    y = pm.Normal('y' , mu=μ , sd=σ , observed=EU50p1d.values)
    trace_g = pm.sample(100000 , tune=1000)
1 Answers

I decided to give your problem a try, but it appears your code has several problems that don't have much to do with the error, while the error doesn't present itself if I follow this procedure:

  • create a new empty project with only a main.py, set up a new Python 3.8 virtual environment for it
  • copy your code into main.py, only removing all unused imports and reformatting
  • create a new test.xlsx with some arbitrary data, change the file reference in the code
  • install the requirements:

Now, you can run the code and it will proceed happily to read the file. Of course, there's some problems in the code that stop it from working:

  • the conversion of EU50p1d.min() to an integer won't work
  • the Uniform() first parameter isn't expected to be a string, but that's what's passed

Also, I don't know what data is in your .xlsx, so of course there may be further problems that don't present themselves. But I doubt those would cause the issues you're describing.

It seems most likely that you're problem starts with not setting up an appropriate virtual environment. Assuming you use a recent Python and just using some example paths which you can change to match needs.

Ensure you're in a location where python on the command line actually executes the Python version you want a virtual environment for, or refer to that python.exe directly:

python -m venv "c:\my_venv\folder"
"c:\my_venv\folder\Scripts\activate"
pip install arviz
pip install pymc3
pip install xlrd

If you're using PyCharm, you can have PyCharm do this for you, just make sure you select the appropriate base Python interpreter before creating the new environment. If you do it manually as described, you'll need to tell PyCharm to use the environment you just created for the project you're working on, in the Project Settings for the Interpreter.

Related