Less features, longer model training time

Viewed 88

I use machine learning algorithm in Malware analysis. When I input some features, I get strange training time. For example:

  1. 4 feature(A,B,C,D), model training time is 3 seconds.
  2. 3 Features(A,B,C), training time is 5 seconds.
  3. 2 features(A, B), training time is 8 seconds.
  4. 1 feature(A), training time is 4 seconds. This kind of result happens on both MLP and Random Forest. In my opinion, the training time should be faster if I use less features, but the result is complete different.

In KNN, the result will be like these:

  1. If I using 6,5,4,3 features(A,B,C,D,E,F), model testing time is about 1.1 seconds, almost the same.
  2. 2 features(A,B), model testing time is 3 seconds.
  3. 1 feature (A), model testing time is 5 seconds.

My dataset has 17K records and using 10-Fold cross-validation. The feature is sort by their entropy, feature A have highest entropy and feature F is lowest. Using Google Colab with sklearn for the testing. I tried several times in different date, and the trend is the same. The feature of my dataset has total 79 features, the appearance only happens with few features.

Thanks for anyone who reply me, I have no idea about it.

1 Answers

It does seem at first glance that having fewer features will result in lower training times. However, depending on which algorithm is being used, this may not be the case. In training, an objective function (loss function) is being minimized by the algorithm. Taking the case of the MLP neural network, if you change the features (especially depending on whether they're informative or not), you're changing the feature space (or "error surface") over which the optimization occurs and possibly the minima of the function will be harder to find, resulting in more steps and longer training in order to satisfy the convergence criteria.

Related