Text Classification with Spacy : going beyond the basics to improve performance

Viewed 1746

I'm trying to train a text categorizer on a training dataset of texts (Reddit posts) with two exclusive classes (1 and 0) regarding a feature of the authors of the posts, and not the posts themselves.
Classes are unbalanced: approximately 75:25, which means that 75% of authors are "0", while 25% are "1".
The whole dataset is composed of 3 columns: the first one representing the author of the post, the second the subreddit the post belongs to, and the third the actual post.

Data

The dataset looks like this:

In [1]: train_data_full.head(5)
Out[1]: 
          author          subreddit            body
0        author1         subreddit1         post1_1 
1        author2         subreddit2         post2_1
2        author3         subreddit2         post3_1
3        author2         subreddit3         post2_2 
4        author5         subreddit4         post5_1

Where postI_J is the J-th post of the I-th author. Notice that in this dataset the same author may appear more than once, if she/he has posted more than once.

In a separate dataset I have the class each author belongs to.

The first thing I did was to group by author:


def proc_subs(l):
    s = set(l)
    return " ".join([st.lower() for st in s])

train_data_full_agg = train_data_full.groupby(["author"], as_index = False).agg({'subreddit':  proc_subs, "body": " ".join}) 

train_data_full_agg.head(5)

Out[2]:
 author                          subreddit                      body     
author1              subreddit1 subreddit2           post1_1 post1_2
author2   subreddit3 subreddit2 subreddit3   post2_1 post2_2 post2_3   
author3              subreddit1 subreddit5           post3_1 post3_2 
author4                         subreddit7                   post4_1
author5              subreddit1 subreddit2           post5_1 post5_2


There are a total of 5000 authors, 4000 used for training and 1000 for validation (roc_auc score). And here is the spaCy code that I'm using (train_texts is the subset of train_data_full_agg.body.tolist() I use for training, while test_texts the one I use for validation).

# Before this line of code there are others (imports and data loading mostly) which i think are irrelevant

train_data  = list(zip(train_texts, train_labels))

nlp = spacy.blank("en")
if 'textcat' not in nlp.pipe_names:
    textcat = nlp.create_pipe("textcat", config={"exclusive_classes": True, "architecture": "ensemble"})
        nlp.add_pipe(textcat, last = True)
else:
    textcat = nlp.get_pipe('textcat')

textcat.add_label("1")
textcat.add_label("0")

def evaluate_roc(nlp,textcat):
    docs = [nlp.tokenizer(tex) for tex in test_texts]
    scores , a = textcat.predict(docs) 
    y_pred = [b[0] for b in scores]
    roc = roc_auc_score(test_labels, y_pred)
    return roc


dec = decaying(0.6 , 0.2, 1e-4)
pipe_exceptions = ['textcat']
    other_pipes = [pipe for pipe in nlp.pipe_names if pipe not in pipe_exceptions]
    with nlp.disable_pipes(*other_pipes): 
        optimizer = nlp.begin_training()
        for epoch in range(10):
        random.shuffle(train_data)
            batches = minibatch(train_data, size = compounding(4., 32., 1.001) )                                                             
            for batch in batches:
                texts1, labels = zip(*batch)
                nlp.update(texts1, labels, sgd=optimizer, losses=losses, drop = next(dec))
            with textcat.model.use_params(optimizer.averages):
                rocs.append(evaluate_roc(nlp, textcat))

Problem

I get poor performance (measured with the ROC on a test dataset i don't have the labels of) also if compared to simpler algorithms one could write using scikit learn (like tfidf, bow, word embeddings etc)

Attempts

I've tried to get better performance with the following procedures:

  1. Various preprocessings/lemmatizations of texts: the best one seems to be removing all punctuation, numbers, stopwords, out of vocabulary words and then lemmatize all remaining words
  2. Tried textcat architectures: ensemble, bow (also with ngram_size and attr parameters): the best seems to be the ensemble, as of spaCy documentation.
  3. Tried to include subreddit information: I did it by training a separate textcat on the same 4000 authors' subreddit column (see point 5. to read how the information coming from this step is used).
  4. Tried to include word embeddings information: using document vectors from en_core_web_lg-2.2.5 spaCy model, I trained over the same 4000 authors' aggregated posts a scikit multi-layer perceptron.(see point 5. to read how the information coming from this step is used).
  5. Then, to mix the information coming from subreddits, posts and document vectors I trained a logistic regression on the 1000 predictions of the three models (i also tried to balance classes in this last step using adasyn

Using this logistic regression, I get ROC = 0.89 over the test dataset. If I remove any of these steps and use an intermediate model, the ROC lowers.

I've also tried the following steps, which again just lowered the ROC:

  1. Use pre-trained models like bert. The code I used is analogous to this one
  2. Tried to balance classes from the beginning, thus using a smaller training set.
  3. Tried to leave punctuation and put a sentencizer at the beginning of the nlp pipeline

Additional information (mainly from comments)

  • Q: What is the ROC of baseline models like Multinomial Naive Bayes or SVM ?
    I have easy access to ROCs evaluated just on texts(no subreddits or vectors). An svm set up like so:
    svm= svm.SVC(C=1.0, kernel='poly', degree=2, gamma='scale', coef0=0.0, shrinking=True, probability=True, tol=0.001, cache_size=200, class_weight=None, max_iter=-1)
    Would give roc (using CountVectorizer bow) = 0.53 (same with rbf kernel, but rbf + class_weight = None or "balanced" gives 0.63 (same without the constraint on cache_size)).Anyway, an XGBregressor with parameters set with gridsearch would give roc = 0.88. The same XGB but also with CountVectorizer subreddits and scikit Doc2Vec vectors (combined with an lr like above) gives about 93. The ensemble code you see above, only on texts would give around 83. With subreddts and vectors (treated as above) gives 89

  • Q: Have you tried not concatenating?
    If i don't concatenate, performance (just on not concatenated texts, so again no vectors/subreddits) is similar to the case in which I concatenate, but I wouldn't then know how to combine multiple predictions for the same author into one prediction. Because remember that I have more comments from each author, and I have to predict a binary feature regarding the authors.

Questions

  1. Do you have any suggestions specifically about the spaCy code that I'm using (e.g. any other way to use subreddits and/or document vectors information)?
  2. How may I improve the overall model?

Any suggestion is highly appreciated.
Please be as explicit as possible in terms of code / explanations / references as I am new to NLP.

0 Answers
Related