I'm trying to train a text categorizer on a training dataset of texts (Reddit posts) with two exclusive classes (1 and 0) regarding a feature of the authors of the posts, and not the posts themselves.
Classes are unbalanced: approximately 75:25, which means that 75% of authors are "0", while 25% are "1".
The whole dataset is composed of 3 columns: the first one representing the author of the post, the second the subreddit the post belongs to, and the third the actual post.
Data
The dataset looks like this:
In [1]: train_data_full.head(5)
Out[1]:
author subreddit body
0 author1 subreddit1 post1_1
1 author2 subreddit2 post2_1
2 author3 subreddit2 post3_1
3 author2 subreddit3 post2_2
4 author5 subreddit4 post5_1
Where postI_J is the J-th post of the I-th author. Notice that in this dataset the same author may appear more than once, if she/he has posted more than once.
In a separate dataset I have the class each author belongs to.
The first thing I did was to group by author:
def proc_subs(l):
s = set(l)
return " ".join([st.lower() for st in s])
train_data_full_agg = train_data_full.groupby(["author"], as_index = False).agg({'subreddit': proc_subs, "body": " ".join})
train_data_full_agg.head(5)
Out[2]:
author subreddit body
author1 subreddit1 subreddit2 post1_1 post1_2
author2 subreddit3 subreddit2 subreddit3 post2_1 post2_2 post2_3
author3 subreddit1 subreddit5 post3_1 post3_2
author4 subreddit7 post4_1
author5 subreddit1 subreddit2 post5_1 post5_2
There are a total of 5000 authors, 4000 used for training and 1000 for validation (roc_auc score).
And here is the spaCy code that I'm using (train_texts is the subset of train_data_full_agg.body.tolist() I use for training, while test_texts the one I use for validation).
# Before this line of code there are others (imports and data loading mostly) which i think are irrelevant
train_data = list(zip(train_texts, train_labels))
nlp = spacy.blank("en")
if 'textcat' not in nlp.pipe_names:
textcat = nlp.create_pipe("textcat", config={"exclusive_classes": True, "architecture": "ensemble"})
nlp.add_pipe(textcat, last = True)
else:
textcat = nlp.get_pipe('textcat')
textcat.add_label("1")
textcat.add_label("0")
def evaluate_roc(nlp,textcat):
docs = [nlp.tokenizer(tex) for tex in test_texts]
scores , a = textcat.predict(docs)
y_pred = [b[0] for b in scores]
roc = roc_auc_score(test_labels, y_pred)
return roc
dec = decaying(0.6 , 0.2, 1e-4)
pipe_exceptions = ['textcat']
other_pipes = [pipe for pipe in nlp.pipe_names if pipe not in pipe_exceptions]
with nlp.disable_pipes(*other_pipes):
optimizer = nlp.begin_training()
for epoch in range(10):
random.shuffle(train_data)
batches = minibatch(train_data, size = compounding(4., 32., 1.001) )
for batch in batches:
texts1, labels = zip(*batch)
nlp.update(texts1, labels, sgd=optimizer, losses=losses, drop = next(dec))
with textcat.model.use_params(optimizer.averages):
rocs.append(evaluate_roc(nlp, textcat))
Problem
I get poor performance (measured with the ROC on a test dataset i don't have the labels of) also if compared to simpler algorithms one could write using scikit learn (like tfidf, bow, word embeddings etc)
Attempts
I've tried to get better performance with the following procedures:
- Various preprocessings/lemmatizations of texts: the best one seems to be removing all punctuation, numbers, stopwords, out of vocabulary words and then lemmatize all remaining words
- Tried textcat architectures: ensemble, bow (also with
ngram_sizeandattrparameters): the best seems to be the ensemble, as of spaCy documentation. - Tried to include subreddit information: I did it by training a separate textcat on the same 4000 authors'
subredditcolumn (see point 5. to read how the information coming from this step is used). - Tried to include word embeddings information: using document vectors from
en_core_web_lg-2.2.5spaCy model, I trained over the same 4000 authors' aggregated posts a scikit multi-layer perceptron.(see point 5. to read how the information coming from this step is used). - Then, to mix the information coming from subreddits, posts and document vectors I trained a logistic regression on the 1000 predictions of the three models (i also tried to balance classes in this last step using adasyn
Using this logistic regression, I get ROC = 0.89 over the test dataset. If I remove any of these steps and use an intermediate model, the ROC lowers.
I've also tried the following steps, which again just lowered the ROC:
- Use pre-trained models like bert. The code I used is analogous to this one
- Tried to balance classes from the beginning, thus using a smaller training set.
- Tried to leave punctuation and put a
sentencizerat the beginning of thenlppipeline
Additional information (mainly from comments)
Q: What is the ROC of baseline models like Multinomial Naive Bayes or SVM ?
I have easy access to ROCs evaluated just on texts(no subreddits or vectors). An svm set up like so:
svm= svm.SVC(C=1.0, kernel='poly', degree=2, gamma='scale', coef0=0.0, shrinking=True, probability=True, tol=0.001, cache_size=200, class_weight=None, max_iter=-1)
Would give roc (using CountVectorizer bow) = 0.53 (same with rbf kernel, but rbf + class_weight = None or "balanced" gives 0.63 (same without the constraint on cache_size)).Anyway, an XGBregressor with parameters set with gridsearch would give roc = 0.88. The same XGB but also with CountVectorizer subreddits and scikit Doc2Vec vectors (combined with an lr like above) gives about 93. The ensemble code you see above, only on texts would give around 83. With subreddts and vectors (treated as above) gives 89Q: Have you tried not concatenating?
If i don't concatenate, performance (just on not concatenated texts, so again no vectors/subreddits) is similar to the case in which I concatenate, but I wouldn't then know how to combine multiple predictions for the same author into one prediction. Because remember that I have more comments from each author, and I have to predict a binary feature regarding the authors.
Questions
- Do you have any suggestions specifically about the spaCy code that I'm using (e.g. any other way to use subreddits and/or document vectors information)?
- How may I improve the overall model?
Any suggestion is highly appreciated.
Please be as explicit as possible in terms of code / explanations / references as I am new to NLP.