Value of k in k nearest neighbor algorithm

Viewed 74346

I have 7 classes that needs to be classified and I have 10 features. Is there a optimal value for k that I need to use in this case or do I have to run the KNN for values of k between 1 and 10 (around 10) and determine the best value with the help of the algorithm itself?

5 Answers

In KNN, finding the value of k is not easy. A small value of k means that noise will have a higher influence on the result and a large value make it computationally expensive.

Data scientists usually choose :

1.An odd number if the number of classes is 2

2.Another simple approach to select k is set k = sqrt(n). where n = number of data points in training data.

Hope this will help you.

You may want to try this out as an approach to running through different k values and visualizing it to help your decision making. I have used this quite a number of times and it gave me the result I wanted:

error_rate = []

for i in range(1,50):
    knn = KNeighborsClassifier(n_neighbors=i)
    knn.fit(X_train, y_train)
    pred = knn.predict(X_test)
    error_rate.append(np.mean(pred != y_test))

plt.figure(figsize=(15,10))
plt.plot(range(1,50),error_rate, marker='o', markersize=9)

There are no pre-defined statistical methods to find the most favourable value of K. Choosing a very small value of K leads to unstable decision boundaries. Value of K can be selected as k = sqrt(n). where n = number of data points in training data Odd number is preferred as K value.

Most of the time below approach is followed in industry. Initialize a random K value and start computing. Derive a plot between error rate and K denoting values in a defined range. Then choose the K value as having a minimum error rate. Derive a plot between accuracy and K denoting values in a defined range. Then choose the K value as having a maximum accuracy. Try to find a trade off value of K between error curve and accuracy curve.

Related