How do I train a sklearn model on a list of numbers?

Viewed 69

I have a sampling type:

Text Target
TEXT Yes
TEXT No
TEXT Yes
... ...

each text can only belong to one class. But the sample contains items with only 2 out of 3 possible target values.

I use the GradientBoostingClassifier model to train text classification, and the .predict_proba function to get a probabilistic answer. But the sample contains only 2 of the 3 possible values, so the function returns answers of type [float,float] (e.g. [0.8,0.2]), although I want an answer of type [float, float, float] (e.g. [0.7,0.2,0.1]). So I converted the sample values as follows:

Text Target
TEXT [1,0,0]
TEXT [0,1,0]
TEXT [1,0,0]
... ...

But the model doesn't want to learn from them. An error is displayed.

~/.local/lib/python3.6/site-packages/sklearn/utils/validation.py in _assert_all_finite(X, allow_nan, msg_dtype)
    108     # for object dtype data, we only check for NaNs (GH-13254)
    109     elif X.dtype == np.dtype('object') and not allow_nan:
--> 110         if _object_dtype_isnan(X).any():
    111             raise ValueError("Input contains NaN")
    112 

AttributeError: 'bool' object has no attribute 'any'

How can I make the model be trained on lists?

My code:

gbc = GradientBoostingClassifier(n_estimators=100, learning_rate=1.0,
                                max_depth=1, random_state=0)
text_clf = Pipeline([('vect', CountVectorizer()),
                     ('tfidf', TfidfTransformer()),
                     ('clf', gbc)
                    ])
text_clf.fit(x,y)
classes = text_clf.classes_
logits = text_clf.predict_proba(X_val)

P.S. I'm going to use not only GradientBoostingClassifier, but also GaussianProcessClassifier, LinearSVC, LogisticRegression, LogisticRegressionCV.

2 Answers

Try to add a dummy line with the missing target value to see if anything changes. This kind of situation could also mean that you don't have enough data for a good machine learning model in general (more data -> more accurate results, no data -> ???)

The Target column should have 3 distinct classes, for example

Text Target
TEXT One
TEXT Two
TEXT Three
TEXT Two
... ...

If the training set contains items with only 2 out of 3 possible target values, then the model will be trained as a 2-class classification model. You need to provide all 3 classes (prefer to be in equal proportion) to the model to train it to be a 3-class classification model.

Related