Background:
Following along with this question when using bert to classify sequences the model uses the "[CLS]" token representing the classification task. According to the paper:
The first token of every sequence is always a special classification token ([CLS]). The final hidden state corresponding to this token is used as the aggregate sequence representation for classification tasks.
Looking at the huggingfaces repo their BertForSequenceClassification utilizes the bert pooler method:
class BertPooler(nn.Module):
def __init__(self, config):
super().__init__()
self.dense = nn.Linear(config.hidden_size, config.hidden_size)
self.activation = nn.Tanh()
def forward(self, hidden_states):
# We "pool" the model by simply taking the hidden state corresponding
# to the first token.
first_token_tensor = hidden_states[:, 0]
pooled_output = self.dense(first_token_tensor)
pooled_output = self.activation(pooled_output)
return pooled_output
We can see they take the first token (CLS) and use this as a representation for the whole sentence. Specifically they perform hidden_states[:, 0] which looks a lot like its taking the first element from each state rather than taking the first tokens hidden state?
My Question:
What I don't understand is how do they encode the information from the entire sentence into this token? Is the CLS token a regular token which has its own embedding vector that "learns" the sentence level representation? Why can't we just use the average of the hidden states (the output of the encoder) and use this to classify?
EDIT: After thinking a little about it: Because we use the CLS tokens hidden state to predict, is the CLS tokens embedding being trained on the task of classification as this is the token being used to classify (thus being the major contributor to the error which gets propagated to its weights?)