Generate 3 circles dataset with three classes in python

Viewed 170

I actually wanted to generate 3 circles one inside the another with three different classes (class0, class1, class2) in python(that is first the bigger circle then inside it the second biggest circle and inside that the third circle). I am only able to generate 2 circles with two classes from the code below. Could anyone please help me with that?

Link of Image for 2 circles

import numpy as np
import pylab as pl
import sklearn.metrics as sm
from sklearn.datasets import make_circles
X, y = make_circles(n_samples=200)
print(X)
print(y)
plt.scatter(X[:,0],X[:,1], marker='o', facecolors='none', edgecolor='r')

2 Answers

Probably there is a more polite way to proceed, but here I show a practical way to do it.

Basically you can generate another two circles with the function make_circles but playing with another hyperparameter, the factor.

What I do, is generating the same main circle, and then a new circle that is multiplied by the factor value (0.6 in my case).

Here is the code:

X, y = make_circles(n_samples=200)
z, w = make_circles(n_samples=200, factor=0.6)

plt.scatter(X[:,0],X[:,1],  facecolors='none', edgecolor='r')
plt.scatter(z[:,0],z[:,1],  facecolors='none', edgecolor='r')

enter image description here

If you want to change the radius of the new circle, play with the factor.

The only problem of this code is that the big circle is duplicated (plotted twice), but, since they are overlapped, there is no visual problem.

According to the documentations of make_circles and make_blobs:

make_circles

sklearn.datasets.make_circles(n_samples=100, *, shuffle=True, noise=None, random_state=None, factor=0.8)
Make a large circle containing a smaller circle in 2d. A simple toy dataset to visualize clustering and classification algorithms.

make_blobs sklearn.datasets.make_blobs(n_samples=100, n_features=2, *, centers=None, cluster_std=1.0, center_box=- 10.0, 10.0, shuffle=True, random_state=None, return_centers=False) Generate isotropic Gaussian blobs for clustering.

You can not make 3 circles using make_circles, but for generating 3 classes, you can use make_blobs.

For having 3 classes circle data, you can use a combination of these two functions as the following example:

import sklearn.datasets as ds
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_blobs

data, labels = ds.make_circles(n_samples=100, 
                               shuffle=True, 
                               noise=0.0, 
                               random_state=42)

center = [[3, 4]]
data2, labels2 = make_blobs(n_samples=100, 
                            cluster_std = 0.2,
                            centers=center,
                            random_state=1)


for i in range(len(center)-1, -1, -1):
    labels2[labels2==0+i] = i+2

print(labels2)
labels = np.concatenate([labels, labels2])
data = data * [1.2, 1.8] + [3, 4]

data = np.concatenate([data, data2], axis=0)

and then the below codes to see the result:

fig, ax = plt.subplots()

colours = ["orange", "blue", "magenta"]
label_name = ["Class1", "Class2", "Class3"]
for label in range(0, len(center)+2):
    ax.scatter(data[labels==label, 0], data[labels==label, 1], 
               c=colours[label], s=40, label=label_name[label])


ax.set(xlabel='X',
       ylabel='Y',
       title='dataset')


ax.legend(loc='upper right')

The results will be as:

enter image description here

Another ways is to use two make_circle functions to generate 4 circles but using 3 of them.

import sklearn.datasets as ds
import numpy as np
import matplotlib.pyplot as plt
from sklearn.datasets import make_blobs

data, labels = ds.make_circles(n_samples=100, 
                               shuffle=True, 
                               noise=0.01, 
                               random_state=42)


data2, labels2 = ds.make_circles(n_samples=100, 
                               shuffle=True, 
                               noise=0.0, 
                               random_state=42)

data2 = data2 * [1.2, 1.8] 

then plot results using:

fig, ax = plt.subplots()
colours = ["orange", "blue"]
label_name = ["Class1", "Class2"]
ax.scatter(data[labels==0, 0], data[labels==0, 1], color='red'
               ,s=40)
ax.scatter(data[labels==1, 0], data[labels==1, 1], color='green'
               ,s=40)
ax.scatter(data2[labels2==0, 0], data2[labels2==0, 1], color='blue',
               s=40)

ax.set(xlabel='X',
       ylabel='Y',
       title='dataset')


ax.legend(loc='upper right')

Then the results are represented as below:

enter image description here

if you want to change the size of circle(3rd circle, here the outer one), you can multiply by different coefficients.

data2 = data2 * [a1, a2] 

where a1 and a2 can be any values but perfectly between 0 and 2. if the values are below 1, the circles will put inner other circles and vice versa.

Related