Correlation coefficient explanation--Feature Selection

Viewed 1693

How to determine the variables to be removed from our model based on the Correlation coefficient .

See below Example of variables:

Top 10 Absolute Correlations:
  Variable 1      Variable 2        Correlation Value
    pdays           pmonths           1.000000
    emp.var.rate    euribor3m         0.970955
    euribor3m       nr.employed       0.942545
    emp.var.rate    nr.employed       0.899818
    previous        pastEmail         0.798017
    emp.var.rate    cons.price.idx    0.763827
    cons.price.idx  euribor3m         0.670844
    contact         cons.price.idx    0.585899
    previous        nr.employed       0.504471
    cons.price.idx  nr.employed       0.490632

correlation matrix heat map of Independent variables":

Below picture is the correlation matrix heat map of Independent variables

Questions:

1)How to remove the one high correlated variable from Correlation-value calculated between two variables

Ex: correlation value between pdays and pmonths is 1.000000 Which variable to be removed from model ?days or pmonths? How the variable is determined ?

2)What is the correlation threshold range considered to drop a variable?ex:>0.65 or >0.90 etc

3)Can you please interpret above Heat map and give your explanation about the variables to be removed and reason for the same?

3 Answers

You could try to use another selection criteria for choosing between each pair of highly-correlated features. For example you can use the Information Gain (IG), which measures how much information a feature gives about the class (i.e., its reduction of entropy [TAL14], [SIL07]). Once you have detected a pair of highly-correlated features (e.g., as you mentioned pdays and pmonths) you can measure the IG of each variable and keep the one with the highest IG. Nevertheless, there are other selection criteria that you could also apply instead of IG (e.g., Mutual Information Maximization [BHS15]).

For the threshold, you can choose the value you want (it depends on your problem). However, for playing safe I would select a high value (e.g., 0.95) although you could also consider those ones around 0.94 or 0.9. Moreover, you can always stablish a high value and then play lowering that value to check the performance of your model.

[TAL14] Jiliang Tang, Salem Alelyani, and Huan Liu. Feature selection for classification: A review, pages 37–64. CRC Press, 1 2014.

[SIL07] Yvan Saeys, Iñaki Inza, and Pedro Larrañaga. A review of feature selection techniques in bioinformatics. bioinformatics, 23(19):2507–2517, 2007.

[BHS15] Mohamed Bennasar, Yulia Hicks, Rossitza Setchi. Feature selection using Joint Mutual Information Maximisation. Expert Systems with Applications, 42(22): 8520- 8532, 2015.

Here I am listing some aspects that are not covered by the other answers

  1. days, months yes, they are correlated - try to generate other features because these are cyclical features so you can generate other features that are not necessarily correlated. read more here

  2. the whole process of feature selection must be done within cross-validation or a hold-out data, otherwise, you are introducing bias and overfitting you model. for example, you can select your features based on a cut-of and then make a model on the rest of data and assess the performance to see how the performance was. you can always use the scikit-learn pipeline to find these thresholds automatically, and nicely. using grid-search.

  3. if you use linear models like lasso - they automatically do these selection!

  4. regarding the threshold, it is hard - but alternatively you can go about top-n features based on number of data points you have available.

  5. correlation-coefficient is a linear model, if covariates are very unlike in linearly they may be very alike in a non-linear space. To learn more you can look into papers related to mutual information to find plenty of examples.

my answers are not directly relevant to the question you are asking - but I am confident they are important to consider and benefit your modeling in a bigger picture.

You can do this easily by using sklearn.

from sklearn.feature_selection import VarianceThreshold

X = [[0, 2, 0, 3], [0, 1, 4, 3], [0, 1, 1, 3]]
selector = VarianceThreshold(threshold=0.2)
X = selector.fit_transform(X)
print(X)

X is the result that removes all unnecessary variable that have low correlation to the others

Related