Minimal pyhf example failing with 'Inequality constraints incompatible'

Viewed 316

I am trying to build a pretty minimal pyhf example: two gaussians, one signal and one background, but I can't get it to work. My python code is:

import pyhf.readxml
import os
from ROOT import TH1F, TFile, TF1

mygaus = TF1("mygaus","TMath::Gaus(x,100,.5)",95, 115)
mygaus2 = TF1("mygaus2","TMath::Gaus(x,110,.2)",95, 115)
mygaus_data = TF1("mygaus_data","TMath::Gaus(x,110,.2)+TMath::Gaus(x,100,.5)",95, 115)

bkg_nominal = TH1F('bkg_nominal', '', 80, 95, 115)
bkg_nominal.FillRandom("mygaus", 10000)

sig_nominal = TH1F('sig_nominal', '', 80, 95, 115)
sig_nominal.FillRandom("mygaus2", 5000)

data_nominal = TH1F('data_nominal', '', 80, 95, 115)
data_nominal.FillRandom("mygaus_data", 10000)

meas = TFile('meas.root', 'RECREATE')
bkg_nominal.Write()
sig_nominal.Write()
data_nominal.Write()
meas.Close()

spec = pyhf.readxml.parse('meas.xml', os.getcwd())
workspace = pyhf.Workspace(spec)

pdf = workspace.model(measurement_name='meas')
data = workspace.data(pdf)
workspace.get_measurement(measurement_name='meas')
best_fit = pyhf.infer.mle.fit(data, pdf)

The XML file, which I basically copied from the example in the documentation, are written like this

meas.xml

<!DOCTYPE Combination  SYSTEM 'HistFactorySchema.dtd'>

<Combination OutputFilePrefix="workspace" >


  <Input>./meas_channel1.xml</Input>

  <Measurement Name="meas" Lumi='1' LumiRelErr='0.1' ExportOnly="False"  >
    <POI>signorm</POI>
  </Measurement>

</Combination>

meas_channel1.xml

<!DOCTYPE Channel  SYSTEM 'HistFactorySchema.dtd'>

  <Channel Name="channel1" InputFile="" >

    <Data HistoName="data_nominal" InputFile="meas.root" />

    <StatErrorConfig RelErrorThreshold="0.05" ConstraintType="Gaussian" />

    <Sample Name="bkg"  HistoName="bkg_nominal"  InputFile="meas.root"  NormalizeByTheory="True" >
      <NormFactor Name="bkgnorm"  Val="1"  High="3"  Low="0"  Const="False"   />
    </Sample>

    <Sample Name="sig"   HistoName="sig_nominal"  InputFile="meas.root"  NormalizeByTheory="True" >
      <NormFactor Name="signorm"  Val="1"  High="3"  Low="0"  Const="False"   />
    </Sample>

  </Channel>

It looks all pretty simple and I am able to plot the histograms. However, when I get this error message:

ERROR:pyhf.optimize.opt_scipy:     fun: nan
     jac: array([nan, nan, nan])
 message: 'Inequality constraints incompatible'
    nfev: 5
     nit: 1
    njev: 1
  status: 4
 success: False
       x: array([1., 1., 1.])
---------------------------------------------------------------------------
AssertionError                            Traceback (most recent call last)
<ipython-input-14-54e7c2f0a645> in <module>
      2 data = workspace.data(pdf)
      3 workspace.get_measurement(measurement_name='meas')
----> 4 best_fit = pyhf.infer.mle.fit(data, pdf)

/usr/local/lib/python3.7/site-packages/pyhf/infer/mle.py in fit(data, pdf, init_pars, par_bounds, **kwargs)
     34     init_pars = init_pars or pdf.config.suggested_init()
     35     par_bounds = par_bounds or pdf.config.suggested_bounds()
---> 36     return opt.minimize(twice_nll, data, pdf, init_pars, par_bounds, **kwargs)
     37 
     38 

/usr/local/lib/python3.7/site-packages/pyhf/optimize/opt_scipy.py in minimize(self, objective, data, pdf, init_pars, par_bounds, fixed_vals, return_fitted_val)
     45         )
     46         try:
---> 47             assert result.success
     48         except AssertionError:
     49             log.error(result)

AssertionError:

which is weird because I don't have any inequality constraint. I think I am doing something dumb, could you please help? Thank you!

2 Answers

Thanks for the good question @robsol90.

If we visually inspect the contents of the model (open the ROOT file and look at the historgrams in TBrowser) or just print out the contents (after converting the XML+ROOT to JSON)

>>> import json
>>> with open("meas.json") as spec_file:
...     spec = json.load(spec_file)
...
>>> print(json.dumps(spec, indent=2, sort_keys=True))

We see that there are many bins with zeros in the model. This is a problem as HistFactory is Poisson based, and as the Poisson p.m.f. is defined strictly for rate parameters greater than 0 these true 0 bins will cause errors (and they do). However, if we simply parse the spec and add a very small offset (epsilon) then the fit is able to proceed without any problem. So this problem actually ends up being very similar to this question (Fit convergence failure in pyhf for small signal model) without it being immediately apparent.

We understand that the toy model you setup was supposed to be minimal and easy, but as in reality you will almost never encounter such a sparse analysis region this toy problem becomes difficult. We will take efforts in the future though to automatically mask bins that are true zeros in the model to avoid this issue for the users altogether.

I'll also give below some code that fixes the problem you have above as well as some additional example code.


First, to be very clear, let's establish our environment

Environment

$ "$(which python3)" --version
Python 3.7.5
$ python3 -m venv "${HOME}/.venvs/question"
$ . "${HOME}/.venvs/question/bin/activate"
(question) $ cat requirements.txt
pyhf[xmlio]~=0.4.0
black
(question) $ python -m pip install -r requirements.txt
(question) $ root-config --version
6.18/04

Code

Let's also break things apart into multiple steps of the code. First let's look at the XML to ROOT code snippet, which I've modified so that there is a more reasonable sampling of the model displayed in the observed data (but did not need to as your original code would work here too).

# XML_to_ROOT.py
from ROOT import TH1F, TFile, TF1


def main():
    left_bound = 95
    right_bound = 115
    n_bins = 80

    # Model makeup
    frac_bkg = 0.95
    frac_sig = round(1.0 - frac_bkg, 2)

    bkg_model = TF1("bkg_model", "TMath::Gaus(x,100,0.5,true)", left_bound, right_bound)
    sig_model = TF1("sig_model", "TMath::Gaus(x,105,0.2,true)", left_bound, right_bound)
    obs_model = TF1(
        "obs_model",
        f"({frac_bkg}*bkg_model)+({frac_sig}*sig_model)",
        left_bound,
        right_bound,
    )

    # Samples from model
    n_sample = 10000
    n_bkg = int(frac_bkg * n_sample)
    n_sig = int(frac_sig * n_sample)

    bkg_nominal = TH1F("bkg_nominal", "", n_bins, left_bound, right_bound)
    bkg_nominal.FillRandom("bkg_model", n_bkg)

    sig_nominal = TH1F("sig_nominal", "", n_bins, left_bound, right_bound)
    sig_nominal.FillRandom("sig_model", n_sig)

    data_nominal = TH1F("data_nominal", "", n_bins, left_bound, right_bound)
    data_nominal.FillRandom("obs_model", n_sample)

    meas = TFile("meas.root", "RECREATE")
    bkg_nominal.Write()
    sig_nominal.Write()
    data_nominal.Write()
    meas.Close()


if __name__ == "__main__":
    main()

Now to make things easier later let's generate our XML and ROOT file and then convert them into a JSON spec

(question) $ python XML_to_ROOT.py
(question) $ pyhf xml2json --output-file meas.json meas.xml

Now, finally, let's adapt the code in your question to make sure that no bins in the model contain true 0s by padding all bins with an offset of 1e-20 (just to demonstrate that the only important thing is that they are non-zero)

# answer.py
import os
import json
import pyhf.readxml
import numpy as np


def main():
    with open("meas.json") as spec_file:
        spec = json.load(spec_file)

    # Pad true zeros to avoid error with evaluating Poisson(x|0)
    epsilon = 1e-20
    bkg = np.asarray(spec["channels"][0]["samples"][0]["data"]) + epsilon
    sig = np.asarray(spec["channels"][0]["samples"][1]["data"]) + epsilon
    spec["channels"][0]["samples"][0]["data"] = bkg.tolist()
    spec["channels"][0]["samples"][1]["data"] = sig.tolist()

    workspace = pyhf.Workspace(spec)

    model = workspace.model(measurement_name="meas")
    data = workspace.data(model)

    best_fit_pars = pyhf.infer.mle.fit(data, model)
    print(f"initialization parameters: {model.config.suggested_init()}")
    print(
        f"best fit parameters:\
        \n * signal strength: {best_fit_pars[0]}\
        \n * nuisance parameters: {best_fit_pars[1:]}"
    )


if __name__ == "__main__":
    main()

Now running we get

(question) $ python answer.py 
initialization parameters: [1.0, 1.0, 1.0]
best fit parameters:        
 * signal strength: 1.000000316044688        
 * nuisance parameters: [0.99884051 1.02202245]

As an extra demonstration that this is really just due to true zeros, consider the following 2 bin example that is engineered to fail with your error.

# fail.py
import os
import json
import pyhf.readxml
import numpy as np


def main():
    with open("meas.json") as spec_file:
        spec = json.load(spec_file)

    # Fails
    bkg = np.asarray([0, 0])
    sig = np.asarray([0, 1])
    obs = np.asarray([1, 1])
    # # Fails
    # bkg = np.asarray([1, 0])
    # sig = np.asarray([0, 0])
    # obs = np.asarray([1, 1])
    # # Fails
    # bkg = np.asarray([0, 0])
    # sig = np.asarray([0, 0])
    # obs = np.asarray([1, 1])
    # # Pass
    # bkg = np.asarray([1e-9, 0])
    # sig = np.asarray([0, 1e-9])
    # obs = np.asarray([1, 1])
    spec["channels"][0]["samples"][0]["data"] = bkg.tolist()
    spec["channels"][0]["samples"][1]["data"] = sig.tolist()
    spec["observations"][0]["data"] = obs.tolist()

    workspace = pyhf.Workspace(spec)

    model = workspace.model(measurement_name="meas")
    data = workspace.data(model)

    best_fit_pars = pyhf.infer.mle.fit(data, model)


if __name__ == "__main__":
    main()

Hi @robsol90 can you dump the full JSON spec pdf.spec and share it here?

Related