PCA on BERT word embeddings

Viewed 2362

I am trying to take a set of sentences that use multiple meanings of the word "duck", and compute the word embeddings of each "duck" using BERT. Each word embedding is a vector of around 780 elements, so I am using PCA to reduce the dimensions to a 2 dimensional point. I expect that words that have the same meaning of "duck" will be clustered together in the graph, but instead there is no recognizable clusters. I am not sure if I am doing something wrong in obtaining the word embeddings or performing PCA on them.

My method of getting the word embeddings:

  tokenized_text = tokenizer.tokenize(marked_text)
  indexed_tokens = tokenizer.convert_tokens_to_ids(tokenized_text)
  segments_ids = [0] * len(tokenized_text)
  tokens_tensor = torch.tensor([indexed_tokens])
  segments_tensors = torch.tensor([segments_ids])
  with torch.no_grad():
    outputs = model(tokens_tensor, token_type_ids=segments_tensors)
    hidden_states = outputs[0]

We are using the last layer of the 12 hidden layers to get the embedding.

For PCA, we're using sklearn.decomposition and calling pca.fit_transform(). Is there a recommended way of normalizing the data (our word embeddings) before calling the function?

0 Answers
Related