Isolation Forest

Viewed 1960

I'm currently working on identifying outliers in my data set using the IsolationForest method in Python, but don't completely understand the example on sklearn:

http://scikit-learn.org/stable/auto_examples/ensemble/plot_isolation_forest.html#sphx-glr-auto-examples-ensemble-plot-isolation-forest-py

Specifically, what is the graph actually showing us? The observations have already been defined as normal/outliers -- so I'm assuming the shade of the contour plot indicates whether that observation is indeed an outlier (e.g., observations with higher anomaly scores lie in darker shaded areas?).

Lastly, how is the following section of code actually being used (specifically the y_pred functions)?

# fit the model
clf = IsolationForest(max_samples=100, random_state=rng)
clf.fit(X_train)
y_pred_train = clf.predict(X_train)
y_pred_test = clf.predict(X_test)
y_pred_outliers = clf.predict(X_outliers) 

I'm guessing it was just provided for completeness in the event someone wants to print the output?

Thanks in advance for the help!

1 Answers
Related